GeoBusinessIQGeoBusinessIQ

Root cause analysis: getting past the plausible explanation to the one you can prove

What this answers

Can we demonstrate this cause produces the fault, rather than merely believing it does?

Every fault attracts an explanation within minutes, and the explanation is usually plausible, comfortable and wrong. Analysis is the work of getting from that first story to something that can be demonstrated, which means evidence, a hypothesis that predicts what you would see, and a test. Without proof, the plant spends money on a change that alters nothing and the fault returns months later to a team who believed it had been solved.

Written for: quality and process engineers, problem-solving teams, production managers.

Describe the problem before explaining it

Most weak investigations start with a cause and search for support. A stronger opening is a precise statement of the fault: which characteristic, on which parts, at which position or cavity, starting when, on which shifts, at what rate, and equally important, where the same fault does not appear. That last question does most of the work. If the defect occurs on one spindle and not the other three, on night shift and not day, or only after a material change, the description has already narrowed the field to a handful of candidates without any theorising.

Ask why repeatedly, but only where you can answer with evidence

Following a chain of consequences backwards is a sound habit and it fails in a predictable way, which is that each step is answered from opinion rather than from observation. The chain then travels quickly to a comfortable conclusion, usually operator error or a training gap, and stops. A useful discipline is to require evidence at each step before continuing, and to treat any answer naming a person rather than a mechanism as a signal to look harder. Human action is generally the last controllable link, not the cause, and a process that relies on nobody making a mistake was never controlled.

Prove it by turning the fault on and off

The strongest available test is to reintroduce the suspected condition deliberately and see whether the defect returns, then remove it and see whether the defect disappears. Where that can be done safely on a machine or a rig, it converts a theory into a demonstration and settles arguments permanently. Where it cannot, a weaker but still useful check is whether the proposed cause accounts for every feature of the description, including the timing and the cases where the fault did not appear. A cause that explains only some of the observations is incomplete.

Two questions, not one: why it happened and why it escaped

Every defect that reaches a customer represents two failures. The process produced it, and the checking system let it through. Investigations that treat only the first leave the detection gap open, so the next unrelated fault escapes exactly the same way. The escape question is frequently the more revealing: the characteristic was never in the control plan, the check existed but at a frequency that missed it, the gauge could not resolve the condition, or the inspector was following an instruction that did not cover the feature. Fixing the detection route protects against faults you have not thought of yet.

Structured methods help, and none of them substitute for evidence

Cause-and-effect diagrams, comparative analysis between good and bad parts, and disciplined team problem-solving formats all improve the odds by forcing breadth before depth and by keeping a record of what was ruled out. They share one weakness: a team can complete any of them beautifully and still be wrong, because the format organises reasoning rather than testing it. Whichever structure is used, the question that determines whether the investigation was worth anything is the same. What did you observe, and what happened when you tried to reproduce it? They earn their place mainly by recording what was considered and rejected, which is what stops the next team repeating the same eliminations from scratch.

Frequently asked questions

How do we know when we have reached the real cause?
When you can make the fault appear and disappear at will by manipulating the condition you identified, or failing that, when the proposed cause accounts for every feature of the problem description including where and when it did not occur. A cause that explains the defects but not their timing, or the affected machine but not the unaffected one, is incomplete. The other practical test is whether the corresponding fix is something you can actually implement and verify.
Is operator error ever a legitimate root cause?
Rarely, and treating it as one usually ends the investigation prematurely. A person made a mistake is the beginning of the useful question, not the answer: why was the mistake possible, what made the correct action unclear or inconvenient, and why did nothing catch it. Processes that depend on sustained human accuracy without any physical or procedural safeguard will produce that mistake again, with a different name attached, regardless of what training follows.
How much time should an investigation take?
Containment is measured in hours and should never wait for analysis. The investigation itself deserves a timescale proportionate to the consequence: a serious escape justifies weeks of engineering attention, a minor recurring irritation does not. What causes damage is the open-ended investigation with no deadline and no owner, which drifts until the problem recurs. Set a date for reporting findings, and if the cause remains unproven by then, say so explicitly rather than adopting the best available guess.

Data limitations

  • Standards are referenced, never reproduced. Pages describe what a standard governs and point to the issuing body; they do not restate its requirements, and conformity is determined by the standard itself and by an accredited assessment, not by anything here.
  • Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.

Explore the graph

Sources

  • NIST Manufacturing Extension Partnership NIST MEP (accessed )
    Covers: A public programme supporting small and medium manufacturers with operational, quality and technology adoption practice.
    Does not cover: Results attributable to any specific manufacturer, or improvement figures transferable to another plant.
    Why it matters: Cited for the operational practice it publishes for smaller manufacturers, not for benchmarks or outcome claims.
    Review cadence: annual
  • International Organization for Standardization ISO (accessed )
    Covers: International standards for quality management, environmental management, occupational health and safety, and industrial processes.
    Does not cover: The content of any standard, conformity decisions, or certification status of any organisation.
    Why it matters: Cited so a reader can reach the issuing body's own public description of a standard. Standard text is never reproduced here.
    Review cadence: annual
  • United Nations Industrial Development Organization UNIDO (accessed )
    Covers: Industrial development analysis, industrial statistics methodology, and manufacturing capability programmes across member states.
    Does not cover: Company-level data, factory costs, supplier information, or real-time production statistics.
    Why it matters: The United Nations agency for industrial development; used for structural framing of how manufacturing sectors develop, never for point figures.
    Review cadence: annual

Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.

Last updated: