Root cause analysis: getting past the plausible explanation to the one you can prove
What this answers
Can we demonstrate this cause produces the fault, rather than merely believing it does?
Every fault attracts an explanation within minutes, and the explanation is usually plausible, comfortable and wrong. Analysis is the work of getting from that first story to something that can be demonstrated, which means evidence, a hypothesis that predicts what you would see, and a test. Without proof, the plant spends money on a change that alters nothing and the fault returns months later to a team who believed it had been solved.
Written for: quality and process engineers, problem-solving teams, production managers.
Describe the problem before explaining it
Most weak investigations start with a cause and search for support. A stronger opening is a precise statement of the fault: which characteristic, on which parts, at which position or cavity, starting when, on which shifts, at what rate, and equally important, where the same fault does not appear. That last question does most of the work. If the defect occurs on one spindle and not the other three, on night shift and not day, or only after a material change, the description has already narrowed the field to a handful of candidates without any theorising.
Ask why repeatedly, but only where you can answer with evidence
Following a chain of consequences backwards is a sound habit and it fails in a predictable way, which is that each step is answered from opinion rather than from observation. The chain then travels quickly to a comfortable conclusion, usually operator error or a training gap, and stops. A useful discipline is to require evidence at each step before continuing, and to treat any answer naming a person rather than a mechanism as a signal to look harder. Human action is generally the last controllable link, not the cause, and a process that relies on nobody making a mistake was never controlled.
Prove it by turning the fault on and off
The strongest available test is to reintroduce the suspected condition deliberately and see whether the defect returns, then remove it and see whether the defect disappears. Where that can be done safely on a machine or a rig, it converts a theory into a demonstration and settles arguments permanently. Where it cannot, a weaker but still useful check is whether the proposed cause accounts for every feature of the description, including the timing and the cases where the fault did not appear. A cause that explains only some of the observations is incomplete.
Two questions, not one: why it happened and why it escaped
Every defect that reaches a customer represents two failures. The process produced it, and the checking system let it through. Investigations that treat only the first leave the detection gap open, so the next unrelated fault escapes exactly the same way. The escape question is frequently the more revealing: the characteristic was never in the control plan, the check existed but at a frequency that missed it, the gauge could not resolve the condition, or the inspector was following an instruction that did not cover the feature. Fixing the detection route protects against faults you have not thought of yet.
Structured methods help, and none of them substitute for evidence
Cause-and-effect diagrams, comparative analysis between good and bad parts, and disciplined team problem-solving formats all improve the odds by forcing breadth before depth and by keeping a record of what was ruled out. They share one weakness: a team can complete any of them beautifully and still be wrong, because the format organises reasoning rather than testing it. Whichever structure is used, the question that determines whether the investigation was worth anything is the same. What did you observe, and what happened when you tried to reproduce it? They earn their place mainly by recording what was considered and rejected, which is what stops the next team repeating the same eliminations from scratch.
Frequently asked questions
- How do we know when we have reached the real cause?
- When you can make the fault appear and disappear at will by manipulating the condition you identified, or failing that, when the proposed cause accounts for every feature of the problem description including where and when it did not occur. A cause that explains the defects but not their timing, or the affected machine but not the unaffected one, is incomplete. The other practical test is whether the corresponding fix is something you can actually implement and verify.
- Is operator error ever a legitimate root cause?
- Rarely, and treating it as one usually ends the investigation prematurely. A person made a mistake is the beginning of the useful question, not the answer: why was the mistake possible, what made the correct action unclear or inconvenient, and why did nothing catch it. Processes that depend on sustained human accuracy without any physical or procedural safeguard will produce that mistake again, with a different name attached, regardless of what training follows.
- How much time should an investigation take?
- Containment is measured in hours and should never wait for analysis. The investigation itself deserves a timescale proportionate to the consequence: a serious escape justifies weeks of engineering attention, a minor recurring irritation does not. What causes damage is the open-ended investigation with no deadline and no owner, which drifts until the problem recurs. Set a date for reporting findings, and if the cause remains unproven by then, say so explicitly rather than adopting the best available guess.
Data limitations
- Standards are referenced, never reproduced. Pages describe what a standard governs and point to the issuing body; they do not restate its requirements, and conformity is determined by the standard itself and by an accredited assessment, not by anything here.
- Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.
Explore the graph
Related manufacturing topics
- Sampling inspection: what a handful of parts can and cannot tell you about a lot
- Skip-lot and reduced inspection: letting lots through on evidence you can defend
- Statistical process control: reading a process while it runs rather than judging it afterwards
- Supplier corrective action requests: raising one, judging the reply, closing it properly
- Supplier quality audits: what a day inside their plant can and cannot tell you
- Supplier quality management: part approval, evidence and what happens after an escape
Across the manufacturing graph
- Yield management: knowing how much good product a process really gives you
- Cycle time: measuring how long the work really takes at each step
- Packaging waste obligations: turning your own packaging into reportable data
- Social audits: being assessed on labour conditions rather than on product quality
- OEE software: settle the definitions before you argue about the figure
- Quoting systems: pricing work you have not done from data you already hold
Sources
- NIST Manufacturing Extension Partnership — NIST MEP (accessed )Covers: A public programme supporting small and medium manufacturers with operational, quality and technology adoption practice.Does not cover: Results attributable to any specific manufacturer, or improvement figures transferable to another plant.Why it matters: Cited for the operational practice it publishes for smaller manufacturers, not for benchmarks or outcome claims.Review cadence: annual
- International Organization for Standardization — ISO (accessed )Covers: International standards for quality management, environmental management, occupational health and safety, and industrial processes.Does not cover: The content of any standard, conformity decisions, or certification status of any organisation.Why it matters: Cited so a reader can reach the issuing body's own public description of a standard. Standard text is never reproduced here.Review cadence: annual
- United Nations Industrial Development Organization — UNIDO (accessed )Covers: Industrial development analysis, industrial statistics methodology, and manufacturing capability programmes across member states.Does not cover: Company-level data, factory costs, supplier information, or real-time production statistics.Why it matters: The United Nations agency for industrial development; used for structural framing of how manufacturing sectors develop, never for point figures.Review cadence: annual
Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.
Last updated: