Machine health analytics: what the technique needs from your data before it works
What this answers
Do we have the signal quality, baseline and failure examples this analysis actually requires, or are we about to generate alerts nobody can act on?
Analysing machine signals to detect developing faults is a genuine technique with hard prerequisites. It needs signals captured fast enough to contain the evidence, a baseline of healthy behaviour under known conditions, and examples of the failures you want detected. Most plants have none of those when they start, which is why so many analytics deployments produce a stream of alerts that maintenance quietly learns to ignore within a few months.
Written for: reliability engineers, data analysts in manufacturing, condition monitoring specialists.
The evidence has to be in the signal you captured
Bearing and gear faults show themselves in vibration at frequencies far above the rate a supervisory system typically samples, and a slowly logged temperature will not reveal them. Motor current carries a surprising amount about mechanical condition, but only if the waveform is captured rather than a smoothed average. Deciding what to capture therefore follows from the failure mode you are trying to see, and that requires knowing which failures actually matter on this machine. Starting from available data instead of from the failure mode produces models built on whatever happened to be logged, which is why they detect so little of interest.
Without a baseline there is nothing to deviate from
Detection works by comparison against normal, and normal must be defined for a specific operating condition. A pump behaves differently at different flows, a machine behaves differently on different products, and a compressor sounds different in summer. A model trained across all of that without knowing which condition applied will treat ordinary operating changes as anomalies and produce alerts on the day the schedule changes. The requirement is unglamorous: capture the operating context alongside the signal, and establish baselines per condition. Establishing them requires a period of running while the equipment is known to be healthy, which is itself an assumption worth checking.
Labelled failures are the resource nobody has
Approaches that learn to recognise a specific developing fault need examples of that fault, with reliable timing of when it began and what it turned out to be. Maintenance records rarely support this: the work order says a bearing was changed, not when the degradation started or whether the bearing was the actual cause. Building a usable labelled history means recording findings carefully from now on, which is a change to maintenance practice rather than an analytics task. In the meantime, anomaly detection against a healthy baseline is more honest, because it flags that something has changed without claiming to know what.
An alert with no action is a nuisance with a technical explanation
The measure of a deployment is not detection performance but what happens after an alert. That requires a defined recipient, an expected response, a way to record what was found, and a rule for what happens when the alert turns out to be nothing. Feeding findings back is the only route to improving thresholds, and it is the step routinely omitted. Also be explicit about false alarms: a system that cries wolf trains its audience to ignore it, and recovering that credibility takes far longer than losing it. Better to alert on fewer things with a clear meaning than on everything a model finds unusual.
Physics beats pattern matching on machines you understand
For well-characterised equipment, established engineering analysis often outperforms a learned model and comes with an explanation. Frequency analysis relates specific vibration components to specific geometry, so a rising component points at a particular element rather than at a general anomaly. Thermodynamic checks on a compressor or pump can identify degradation from measurements you already collect. These methods need fewer examples, survive changes in operating condition better, and produce a diagnosis an engineer can argue with. Reserve statistical learning for complex assets where no adequate physical model exists, and for combining evidence across several sources.
Frequently asked questions
- Why do our analytics alerts get ignored by maintenance?
- Almost always because too many of them were wrong, or because the alert did not say what to do. Credibility is spent quickly and rebuilt slowly. The remedies are to raise the threshold so that alerts are rare and meaningful, to route them to a named person with an expected response, and to close the loop by recording what was actually found. Feeding those findings back is what turns a noisy detector into something the shift team consults voluntarily.
- How much healthy running do we need before a model is useful?
- Enough to cover the operating conditions the machine genuinely experiences, including different products, loads and seasonal ambient conditions. A baseline collected during one product campaign will flag the next campaign as abnormal. There is no single figure, because it depends on how varied the duty is: a machine running one product continuously needs far less than one switching between many. Record the operating context alongside the signal so baselines can be built per condition rather than blended into an average.
- Do we need machine learning to detect developing faults?
- Often not. For rotating equipment, established vibration analysis relates specific frequency components to specific defects and produces a diagnosis rather than an anomaly score. Simple trending of temperature, current draw or efficiency catches a surprising amount. Statistical learning earns its place on complex assets with no adequate physical model, or where evidence from several sources needs combining. Starting with the engineering methods also builds the signal quality and baselines that any later model would have required anyway.
Data limitations
- Plant, process, utility and equipment material is business intelligence, not engineering design. Layout, structural, electrical, mechanical, pressure, ventilation and fire-safety decisions require a qualified engineer working to the codes in force at the site.
- Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.
Explore the graph
Related manufacturing topics
- Machine tending automation: buffers, part presentation and the machine interface decide the cell
- Machine vision: lighting, optics and why a camera sees less than you think
- Machinery safety functions: what a machine has to do when something goes wrong
- Marking and reading parts in production: where identification actually breaks
- OT and IT convergence: two professions with opposite instincts sharing one set of networks
- Packaging line automation: the stoppages come from the materials, not the machinery
Across the manufacturing graph
- PLM to ERP: getting the engineering definition into the system that buys and builds
- Serialisation systems: allocating, applying and accounting for unit identity
- Equipment replacement: choosing between keeping, rebuilding and replacing a machine
- Line balancing: sharing work content so no station sets the pace alone
- Factory design: writing the brief the building has to satisfy
- Fixed-position layout: when the product is too big to move
Calculators
Logistics & supply chain
Sources
- National Institute of Standards and Technology — NIST (accessed )Covers: Measurement science, manufacturing technology research, cybersecurity frameworks, and industrial standards support.Does not cover: Certification of products, endorsement of vendors, or costs for any specific implementation.Why it matters: A United States federal research institute whose public material covers measurement, manufacturing technology and control-system security.Review cadence: annual
- International Electrotechnical Commission — IEC (accessed )Covers: International standards for electrical, electronic and related technologies, including industrial automation and machinery safety.Does not cover: Standard text, conformity decisions, or product approval.Why it matters: Cited for the origin of electrotechnical and automation standards referenced on automation and machinery pages.Review cadence: annual
- United Nations Industrial Development Organization — UNIDO (accessed )Covers: Industrial development analysis, industrial statistics methodology, and manufacturing capability programmes across member states.Does not cover: Company-level data, factory costs, supplier information, or real-time production statistics.Why it matters: The United Nations agency for industrial development; used for structural framing of how manufacturing sectors develop, never for point figures.Review cadence: annual
Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.
Last updated: