GeoBusinessIQGeoBusinessIQ

Field failure analysis: getting the broken part back and reading it honestly

What this answers

What does this returned part tell us about how it failed, and does that match anything in our production record?

A part that failed in service carries information no test rig will give you, because it was loaded, installed and abused the way real products are. Most of that information is destroyed before an engineer sees it — by the person who removed it, the packaging it travelled in, the cleaning somebody did to be helpful, or the months it spent in a box. Recovering it starts long before the teardown.

Written for: reliability engineers, quality engineers, service and technical support managers.

The evidence degrades from the moment of removal

Whoever takes the part off is your only witness to its as-failed condition, and they are usually focused on restoring service. Give them a short, blunt request: photograph it in place before removal, do not clean it, note what the equipment was doing when it failed, record hours or distance, keep any fragments, and bag it dry. Fracture surfaces corrode, contamination washes off, wear patterns are polished by handling, and electrical damage looks different after somebody has probed it. A brief field form completed at the point of removal is worth more than a long questionnaire completed later from memory.

Teardown is destructive, so sequence it deliberately

Each step of an examination removes evidence for the next one, which is why an unplanned strip-down by a curious engineer is the most expensive mistake in the process. Agree the sequence first: external inspection and photography, dimensional and functional checks where possible, then non-destructive examination, then sectioning, then anything metallurgical or chemical. Record at every stage rather than at the end. Where the failure may become a legal matter, custody has to be documented and the part may need to be examined jointly, which is a decision to take before the first cut rather than after.

No fault found is a finding, not a dead end

A high proportion of returns test as functional on the bench. Treating that as the answer wastes the return and irritates the customer, who watched the thing fail. The usual explanations are worth working through: the failure is intermittent and appears only under temperature, vibration or load the bench does not apply; the fault lies in the installation, the mating part or the system rather than the component; the diagnostic procedure in the field points at the wrong item; or the symptom is a performance expectation rather than a defect. Each of those has a different owner and a different fix.

Connecting the part back to how it was made

The analysis becomes actionable when the physical evidence meets the production record. Read the marking, identify the lot, and pull what exists for it: material certification, process parameters, inspection results, tooling in use, whether a deviation was in force, who ran it. Then ask the harder question of whether the failed population clusters — one lot, one machine, one shift, one supplier delivery, or spread evenly across everything you made. A cluster points at a manufacturing cause; an even spread across a long production history points at the design or the application instead.

Making the loop close, not just the report

Field analysis earns its cost only when findings reach the people who can act: the designer who set the margin, the process owner whose control did not catch it, the technical author whose installation instruction was ambiguous, and whoever writes the risk analysis for the next product. That requires a standing route from the analysis to those functions, and a check that the failure mode discovered was actually added to the design and process risk documents. Reports that circulate within engineering and never touch the control plan are the reason the same mechanism reappears on the successor product.

Frequently asked questions

How do we get customers to return failed parts at all?
Make it easy and make it worth their while. Pre-paid packaging, a simple form, a named contact, and a commitment to tell them what was found all raise return rates substantially. Where the part is replaced under warranty, tie the credit to the return so the commercial process does the collection for you. Expect returns to be biased towards accessible, high-value items, and treat the sample as informative rather than representative.
Should we send failed parts to an external laboratory?
For metallurgical fracture work, chemical analysis, electronics failure localisation and anything likely to be contested, yes — the equipment and the independence are both worth paying for. Keep the initial visual and functional work in-house, because it decides what to send and what question to ask, and a laboratory given a part with no context will answer a general question expensively. Agree the scope, the retention of the sample and the reporting format before shipping it.
How many field returns do we need before we act?
That depends entirely on severity. A single failure with a safety consequence justifies immediate containment while the analysis runs. A cosmetic complaint may need a trend across many units before a change is worth its disruption. What should not vary is the threshold for looking: every return with a plausible manufacturing cause deserves at least an examination and a note against the lot, because the second occurrence is only recognisable if the first was recorded.

Data limitations

  • Standards are referenced, never reproduced. Pages describe what a standard governs and point to the issuing body; they do not restate its requirements, and conformity is determined by the standard itself and by an accredited assessment, not by anything here.
  • Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.

Explore the graph

Sources

  • National Institute of Standards and Technology NIST (accessed )
    Covers: Measurement science, manufacturing technology research, cybersecurity frameworks, and industrial standards support.
    Does not cover: Certification of products, endorsement of vendors, or costs for any specific implementation.
    Why it matters: A United States federal research institute whose public material covers measurement, manufacturing technology and control-system security.
    Review cadence: annual
  • United Nations Industrial Development Organization UNIDO (accessed )
    Covers: Industrial development analysis, industrial statistics methodology, and manufacturing capability programmes across member states.
    Does not cover: Company-level data, factory costs, supplier information, or real-time production statistics.
    Why it matters: The United Nations agency for industrial development; used for structural framing of how manufacturing sectors develop, never for point figures.
    Review cadence: annual

Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.

Last updated: