Root Cause Analysis
Root cause analysis is the systematic examination of a problem that treats its origin as the point of correction. It separates three things everyday conversation runs together: the symptom, which is observed and measured, the chain of causes leading to it and the root cause, the system-level condition whose removal stops the recurrence and which the organisation can actually correct. Two techniques sit under the name, the fishbone diagram and the Five Whys, and they answer two different questions: where could this effect come from and why did this particular link occur. As the entry point to the family, it carries what the two have in common, the problem statement, the evidence, the verification of a candidate cause and the hand-off to corrective action, then routes to whichever suits the shape of the causality.
Goal
Root cause analysis establishes where a problem comes from, so that the correction lands on its origin. Choosing the remedy comes after it and belongs to other techniques.
The deliverable is a set of verified causes, written as causal statements, each tied to the evidence that made it stand, then handed over to corrective action. Rarely a single cause: BABOK runs the analysis iteratively because several root causes can contribute to the same effect, and the PMBOK Guide notes that one root cause may sit behind several variances or defects. A report that returns a single cause for an operational incident has usually stopped the analysis too early.
Use is reactive or proactive. In reactive analysis, a problem has occurred and you look for where to correct it. In proactive analysis, you start from a weak signal that already exists, a near miss or a measured drift and work back to the condition producing it, before the failure arrives. The raw material is still an observed fact. BABOK names both uses, with a qualification that comes from the standards bodies: IEC 62740 restricts its own scope to a posteriori analyses, while allowing that some techniques adapt to design work and risk assessment.
Four activities structure the work, whichever technique is chosen: defining the problem statement, collecting data on the nature of the effect, its magnitude, its location and its timing, identifying the cause and identifying the corrective action that will prevent or limit recurrence. The two techniques of the family occupy the third activity. The other three run the same way in both cases, and the discipline they impose weighs more on the result than the choice of technique does.
Usage
When to use it
- Problem recurring despite successive fixes: the previous correction landed on the effect, the mechanism is still intact.
- Measured gap between expected and actual performance: magnitude, location and period are known, the problem statement can be written today.
- Before specifying a requested solution: requirements then attack the cause of the problem.
- Current-state analysis: establishing that a problem asserted by the organisation is real and where it originates.
- After an incident or a failed release: a corrective action has to be chosen and then defended to governance.
- Fragile area where signals already exist: near misses, measured drifts, the analysis works back to the condition producing them.
When not to use it
- Failure mode to be imagined on a system that has not run yet: nothing to establish after the fact, use failure mode and effects analysis (FMEA) or risk analysis.
- Cause already established and evidenced: only the choice among remedies is left, move to decision analysis or multi-criteria decision analysis.
- Deliberate or blameworthy act: the approach assigns neither responsibility nor liability, the route is administrative or through human resources.
Symptom, cause, root cause
Three terms run together in ordinary usage and everything else depends on keeping them apart. The symptom is what is observed and measured, the visible outcome that made someone open the analysis. OSHA puts it plainly: correcting only the immediate cause removes a symptom of the problem, not the problem. A cause is a circumstance or set of circumstances that leads to failure or success, in IEC 62740's definition; between the origin and the effect runs a chain of circumstances whose last member is the immediate cause. The root cause is the underlying, system-related condition whose removal stops the recurrence and which identifies a correctable failure of the working system.
Three questions make those definitions usable in the room. If this condition were removed, would the event stop recurring? That is the PMBOK Guide's non-recurrence criterion. Is it within the organisation's power to remedy? A condition outside it remains a constraint to design around, and ASQ counts that controllability test among the three a root cause has to meet. Does an action follow that someone can be assigned? A cause that yields no workable recommendation has not finished being analysed.
OSHA's published illustration is still the sharpest. A worker slips on spilled oil, the investigation concludes "oil on the floor", the remedy is to mop it up and advise more care. The root cause was the absence of a mechanical integrity programme that would have caught the leak. Both analyses cover the same incident and only one prevents the next.
The discipline around the analysis
The problem statement
It carries the four dimensions of data collection: the nature of the effect, its magnitude, its location and its period. "Of the 500 applications processed last quarter, 160 exceeded the contractual 30-day turnaround" is a statement. "Processing is too slow" is a complaint: it produces a vague analysis. Two drafting faults distort the session before it starts: a statement containing a cause, "the validation rules are wrong", has already decided the outcome; a statement containing a solution, "we are missing a second approval step", is a remedy in disguise and leaves nothing to analyse.
Facts before analysis
IEC 62740 sets out the process in five ordered steps: initiation, establishing the facts, analysis, validation, presentation of results. Evidence is gathered before the causal work. Gathered afterwards, it only serves to support the conclusion the room reached in the first ten minutes. An organisation that opens its analyses with a round of hypotheses has inverted the first two steps, and it will spend the session defending the first explanation anyone voiced.
Team composition
The standard gives the analysis team a clause of its own, and the VA guide is more concrete: frontline people are usually best placed to identify problems and solutions. Every analysis runs into a knowledge ceiling, the ceiling of what the room knows about the domain. Bringing in the person who runs the failing process, the person who owns the system involved and one outside pair of eyes costs three invitations and moves that ceiling.
Verification
A candidate cause is a hypothesis until evidence separates it from the others. IEC 62740 makes this a validation clause of the process in its own right. BABOK carries the same requirement inside the fishbone diagram, stating that the group has identified only potential causes and that further analysis, ideally data-based, is still needed to validate the actual cause. Skipping this step is the most expensive error in the whole family, because the deliverable looks finished: a wall of plausible causes, a chain that hangs together end to end and no data behind any of it.
Verifying has a concrete shape, the same whichever technique was used. A candidate cause makes a prediction, so the test is to look where the effect should also appear and where it should not: another site, another period, another team, another channel. If the proposed cause is real, the gap shows up wherever the condition is present and nowhere else. The evidence is taken from a system: a count, a date, a log entry or an extract. When the data does not exist, the honest move is to fund a targeted measurement over a few weeks rather than to settle it by show of hands: an analysis that leaves the room marked "to be measured" with an owner beats a root cause voted through.
Writing the cause down
A cause is written in a fixed cadence, the VA guide's: the cause, then "led to", then the effect, then "which increased the likelihood that", then the event. Five rules of causation govern that drafting and they hold well beyond patient safety. Show the cause-and-effect relationship explicitly, where an isolated observation of the "a staff member was fatigued" kind stops at the effect. Use precise descriptors: "the poorly written manual" does not name what the manual lacked. A human error must have a preceding cause. A procedure violation likewise requires a cause that precedes it. An omission is causal only where a duty to act already existed. The last three are the general form of a discipline each of the two techniques applies on its own side: an analysis never ends on a person.
Neither blame nor liability
IEC 62740 excludes the assignment of responsibility or liability from its scope. The IHI says the same thing from the other end: the review does not address individual performance, and its process is not recommended for blameworthy events, criminal acts or deliberately unsafe acts, which belong to the administrative and human-resources routes. The reason is as practical as it is ethical. An organisation that has watched an analysis end in a sanction will never supply useful evidence again, and every later analysis depends on that supply.
Choosing between the two techniques
Choosing the technique is a step of the analysis, argued like the others. IEC 62740 devotes a full clause to it and offers attributes for comparing techniques against each other; Doggett builds a selection framework of his own on documented characteristics, the ability to surface causal interdependencies, to keep a group focused, to stay readable once complete. The deciding question fits on one line: is a single causal chain plausible and known, or does the field have to be swept first?
Eight criteria separate the two techniques, drawn from what each one produces and from the way each one fails.
| Criterion | Fishbone diagram | Five Whys |
|---|---|---|
| Shape of the causality | Several contributing factors, plausibly in parallel. | A single linear chain is plausible. |
| Direction of the question | Breadth: where could this effect come from? | Depth: why did this link occur? |
| Knowledge of the field | The families of causes are not known in advance, the room risks missing a whole class. | One failure mode is plausible and the room already knows where to look. |
| Working setting | A facilitated group, a wall or a shared board, categories agreed before the session. | A small group or a plain conversation, with no materials. |
| State of the evidence | Data will be sought after the session, to adjudicate between candidates. | Evidence is available link by link, in the room or right afterwards. |
| Facilitation cost | A prepared session, plus a validation phase to be funded and owned. | A few dozen minutes, one facilitator. |
| Trace left behind | A map of everything that was considered, useful when the sweep has to be shown. | A single chain, thin as an audit record on its own. |
| Output produced | A ranked set of candidate causes, to be validated against data. | A chain ending on a correctable condition. |
BABOK describes their composition itself: the Five Whys are used alone or inside the diagram, once all the ideas are captured on the figure, to drill down to the root causes. The junction is a candidate the data has confirmed. The fishbone diagram says where to dig, the Five Whys dig. The routing rule fits in one sentence: sweep first when you do not know the field, dig first when you do, do both in that order when the decision commits the organisation.
Each fails in its own way, and the choice is also a choice of which failure you would rather watch for. A diagram that looks complete gets taken for an answer: the room admires the sweep, no branch is dug, validation never happens. A single chain gets imposed on a multicausal reality, lets itself be built backwards by whoever already knows the conclusion and ends comfortably on human error. Both are watched for during the session: on later reading, a complete diagram and a coherent chain look exactly like successful work.
The field is wider than these two techniques. IEC 62740 describes a dozen of them in Annex C, among them events and causal factors charting, fault tree analysis and Tripod Beta, alongside four causal models including Reason's Swiss cheese model. A few neighbouring methods are regularly confused with root cause analysis. Pareto analysis selects the problem that deserves the analysis and therefore sits ahead of it. A3 is a reporting and problem-solving format: it hosts a cause analysis without supplying the method. 8D and the Kepner-Tregoe method are complete structured approaches in which causal analysis is only one phase.
One wider boundary is worth drawing: there are situations where the whole family is the wrong instrument. A process that has always produced this level of result, with variation that stays inside its usual spread, presents no event to explain. Its variation is common, that is, produced by the process itself, as against special variation, which signals that an outside factor intervened at an identifiable moment. Opening an analysis on common-cause variation manufactures a cause, because a group asked to find one always finds one, and the correction that follows will address a month that was in no way remarkable. The prior move is therefore to separate the two, by plotting the measure on a control chart or simply by comparing the period under suspicion with the full history. A point outside the usual limits opens the analysis; a month slightly worse than the others calls for work on process capability, which is a different discipline.
From cause to action
The analysis stops at the verified cause. The fourth activity, identifying the action, defines the correction that will prevent recurrence, and then the work passes to those who decide and execute. That hand-off is the documented breaking point of the practice: the IHI renamed the process RCA², Root Cause Analyses and Actions, to mark that prevention requires actions. Peerally and his co-authors count poor follow-through on actions among the three main defects of root cause analysis as it is practised.
The quality of an action is judged on its strength, meaning on what it depends on to hold. The hierarchy established by the VA and published by the IHI sorts actions into three levels.
| Level | What the action changes | Common forms |
|---|---|---|
| Strong | The failure becomes impossible or very hard, independently of the care taken by whoever does the work. | Blocking control, field made mandatory in the form, removal of the manual step, automated transfer between two systems. |
| Intermediate | The failure stays possible, its likelihood drops markedly. | Checklist, contextual alert at the point of entry, redundancy on a critical step, added staffing on a bottleneck. |
| Weak | The outcome depends entirely on people's memory and vigilance. | Training, reminder of the rule, update to a directive, manual double check. |
"Train the teams and update the directive" as a complete action plan is the commonest way for a correct analysis to produce nothing. Each retained action carries a named owner, a due date and a measure that will say whether it worked, and the plan goes to leadership for approval and funding. For a certified organisation, ISO 9001 places this move inside the management system: faced with a nonconformity, the organisation evaluates the need for action to eliminate the cause or causes, so that it does not recur or occur elsewhere.
AI considerations
Three uses hold up, all of them on the material of the analysis. Sweeping the past brings the problem statement together with the incident history, the tickets and the logs, pulls out candidates nobody would have phrased and attacks the room's knowledge ceiling. Upstream triage clusters thousands of incidents by similarity and shows which problem deserves an analysis, a selection nobody performs by hand. Write-up drafts the causal statements in the expected cadence from the session notes, then reads the action plan against the action hierarchy and flags the one that contains nothing but training and a directive.
The limit falls exactly on the two moves that make the analysis. A model proposes candidates and reviews an action plan against the hierarchy, two operations on text. It can do neither of the two things the result depends on: verifying a candidate against evidence, which means going into a system for the count, the date or the log entry and comparing periods and sites, and choosing the technique, which assumes knowing what the room understands of the domain, what data it holds and what trace the organisation will have to show. This whole family of techniques fails the same way, by accepting an explanation because it is plausible, and generated text is plausible by construction: a proposal from a model therefore enters the analysis as a hypothesis to be tested. The materials, finally, are often personal or sensitive, access logs, client files, human-resources records: processing them in an external tool assumes the legal basis and the authorisation that data protection requires.
Visualisations
What an analysis leaves behind is a written file. It has four pieces, three of them made of rows and columns.
| Piece of the file | The columns it carries | What it allows once the analysis is closed |
|---|---|---|
| Problem statement | The measured effect, its magnitude, its location, its period. | Reopening the file next quarter on the same measure and saying whether the problem has receded. |
| Register of retained causes | One row per cause: the causal statement, the evidence that made it stand, the reason for not pushing further. | Replaying the reasoning without having been in the room and contesting it cause by cause. |
| Action plan | One row per action: its level, its owner, its due date, its control measure. | Tracking execution and spotting at a glance a plan that contains only weak actions. |
A table held in a document sorts, reads back by column and reopens next quarter. The fourth piece is the figure produced by the technique chosen, and its form belongs to that technique. What remains is the decision that comes before everything, the choice of technique: it is drawn, because it is made of questions and branches.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | Medium | The problem statement needs a measure, a scope and a period, so a first pass on the data. Add to that the composition of the team, which decides the quality of the result, and the argued choice of technique. |
| Execution | Medium | The analysis session runs from thirty minutes to half a day depending on the technique chosen. The real cost lies elsewhere, in the validation phase that follows: going after the data that adjudicates between candidate causes is the part schedules forget to fund. |
| Documentation | Medium | The file runs to a few pages, statement, causes, evidence, actions. The lasting load is tracking the action plan through to the measurement of its effect, without which the analysis produces nothing. |
Tooling
The session itself needs nothing: a whiteboard, a flip chart or a shared document kept live will do, and the analysis is run where the problem occurs. Remotely, a digital whiteboard (Miro, Mural) replaces the wall with nothing lost, provided the content is then copied into a document, which becomes the register.
The spreadsheet is the tool of the cause register and the action plan, because the evidence column becomes a visible constraint there. The data that adjudicates between candidates comes from the systems: a query on the production database, an extract from the ticketing tool, a management dashboard. This is the most decisive tooling item of all, since it decides whether validation happens.
Two families of specialised software come in downstream. Incident management platforms offer post-mortem templates where the causal analysis is a structured field and where actions become tracked tickets, which addresses the follow-through defect directly. Quality management software (TapRooT, Intelex, Cority) imposes a template, keeps an auditable trail and allows analyses to be compared with one another, which matters as soon as they accumulate. None of these tools asks a question, verifies evidence or picks a strong action.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.40 Root Cause Analysis: the definition, reactive and proactive analysis, the four activities, the two techniques of the family and their composition.
- PMI, A Guide to the Project Management Body of Knowledge (PMBOK Guide), 8th edition: pmi.org/standards/pmbok. The definition of the root cause, the non-recurrence criterion and the fact that one root cause may sit behind several variances.
- IEC 62740:2015, Root cause analysis (RCA): webstore.iec.ch/en/publication/21810. The five-step process and its validation clause, the definition of a cause, the restriction to a posteriori analyses, the exclusion of responsibility and liability, the selection of techniques and the inventory in Annexes B and C.
- OSHA and EPA, The Importance of Root Cause Analysis During Incident Investigation (OSHA 3895): osha.gov/publications/OSHA3895.pdf. The definition of the root cause, the contrast between symptom and immediate cause and the spilled-oil illustration.
- American Society for Quality, What is Root Cause Analysis?: asq.org/quality-resources/root-cause-analysis. The three criteria a root cause has to meet, including controllability by the organisation, after Rooney and Vanden Heuvel, Root Cause Analysis for Beginners, Quality Progress, July 2004.
- VA National Center for Patient Safety, Guide to Performing a Root Cause Analysis, rev. 02/2021: patientsafety.va.gov/docs/RCA-Guidebook_02052021.pdf. The cadence of the causal statement, the five rules of causation and the place of frontline people on the team.
- Institute for Healthcare Improvement, RCA²: Improving Root Cause Analyses and Actions to Prevent Harm: ihi.org/library/tools/rca2. The renaming of the process around actions and the exclusion of blameworthy events.
- Institute for Healthcare Improvement, Patient Safety Essentials Toolkit: Action Hierarchy: ihi.org/SafetyToolkit_ActionHierarchy.pdf. The three action levels and the rule of at least one strong or intermediate action per cause.
- A. M. Doggett, Root Cause Analysis: A Framework for Tool Selection, Quality Management Journal 12(4), 2005: doi.org/10.1080/10686967.2005.11919269. The framework for selecting a causal analysis tool on documented performance characteristics.
- M. F. Peerally, S. Carr, J. Waring and M. Dixon-Woods, The problem with root cause analysis, BMJ Quality & Safety 2017;26:417-422: qualitysafety.bmj.com/content/26/5/417. The focus on a single cause, the weakness of the analyses and poor follow-through on actions.
- ISO 9001:2015, §10.2 Nonconformity and corrective action: iso.org/standard/62085.html. The requirement to evaluate the need for action to eliminate the cause or causes of a nonconformity.

