This isn't an argument against uptime tracking — if you're measuring OEE, you're already ahead of most plants. It's an argument about what uptime data can and cannot do, and about the question every downtime report leaves hanging: fine — but why was the machine off?
The Scoreboard Problem
Machine monitoring platforms and OEE dashboards have become genuinely good at one thing: counting. Connect a sensor or a PLC signal, and you get run hours, stop hours, availability percentages, a downtime Pareto by machine and by shift. For many SMEs this is the first time the losses have ever been visible in one place, and the number is usually a shock.
But visibility is where this layer ends. An availability figure is an accounting entry — it records that value was lost, after it was lost. It's a scoreboard: essential for knowing whether you're winning, useless for telling you how to win the next match. Across UK manufacturing the scoreboard has been showing the same result for years — unplanned downtime still costs up to £736 million every week. [1]
If measurement alone fixed downtime, the plants with the prettiest dashboards would have the highest availability. They don't, reliably — because the dashboard reports the cost of the problem while saying nothing about its cause. Data-rich, diagnosis-poor: the same trap we've written about in why engineering teams are drowning in data.
The Ladder: That → Where → Why → Never Again
Understanding a machine failure is a ladder of four questions, and each rung narrows the one above it. Most plants stop climbing after the first — some after the second. The value is at the bottom.
| Layer | Question it answers | Where it comes from | Its limit |
|---|---|---|---|
| Uptime / OEE | That you lost — and how much | Run signals, PLC counts, timers | Counts the loss after it happened |
| Reason codes | Where it went — which bucket | Operator entry at the machine | “Breakdown” is a label, not a reason — and it's opinion, after the event |
| Condition monitoring | Why, physically | Vibration, temperature, current, oil — measured physics | Explains and predicts the failure — but someone still has to act |
| The fix | How it never happens again | Engineers on site — root-cause, engineer out, verify | The step the software market hands back to you |
Each rung is only as useful as the rung below it. An availability figure with no reasons is a number you can worry about but not act on. A reason code of “breakdown” with no physics is a bucket you can chart but not empty. And a diagnosis with no fix is a better-informed way of waiting for the same failure — which is the whole argument of closed-loop reliability.
“Breakdown” Is a Label, Not a Reason
The standard answer to the “why” question is downtime reason codes — the operator picks a category when the line stops. And for the organisational half of downtime, they genuinely work: changeover, no material, no operator, waiting on quality are process problems, visible to the person standing at the machine, and fixable with planning and scheduling.
Then there's the other bucket. Every downtime Pareto has a bar called unplanned breakdown — usually the tallest, always the most expensive per hour — and it is precisely the bucket where reason codes stop working. Because “breakdown” isn't a reason. It's a location where a reason should be. The reason is a bearing that had been degrading for three weeks, a coupling knocked out of alignment at the last rebuild, oil quietly filling with wear metal. None of that is visible to an operator at the HMI, and none of it fits in a dropdown.
- Reason codes are opinion, recorded after the event — a busy operator's best guess, tagged while trying to restart the line. Miscoding and the giant “other” bucket are universal.
- Condition data is measurement, recorded before and during it — vibration signatures, temperature trends, motor current, fluid analysis. Physics doesn't misremember what happened.
The tallest bar on your downtime Pareto is the one nobody can explain. That bar is the case for condition monitoring, drawn by your own uptime data.
Past Tense vs Future Tense
There's a second difference between the layers, and it matters even more than the why: tense. Uptime data is written in the past tense — a downtime event has to exist before it can be counted. By the time the dashboard knows anything, you've already paid for the knowledge in lost production.
Condition data is the only layer that speaks in the future tense. Most mechanical failures develop over weeks or months, and they announce themselves — rising vibration at a fault frequency, creeping temperature, wear metals climbing in the oil — long before the machine actually stops. That's the entire premise of condition monitoring as ISO 17359 frames it: measure the health continuously, detect the departure from normal, and intervene while the machine is still running. [3] It's the difference between a record of your downtime and a warning of it. The reliability theory behind that warning window is the P-F curve.
Put bluntly: uptime monitoring measures downtime; condition monitoring prevents it. Deloitte's figures for programmes that act on condition data — 10–20% more uptime, 5–10% lower maintenance cost — are the size of the gap between the two tenses. [5]
Not Rivals — a Stack
None of this means ripping out your machine monitoring platform. The layers do different jobs, and the top of the ladder is where the bottom gets its priorities from. Availability is one of OEE's three factors [2], and your OEE data is the cheapest, fastest way to answer the question every condition-monitoring project starts with: which machine first?
- Uptime data sizes the prize — it tells you the breakdown bar costs you 60 hours a quarter, and it names the machine at the top of the Pareto. If you want that in pounds, our downtime cost calculator does the arithmetic.
- Condition monitoring explains and predicts — on that named machine, it converts “breakdown: 14 hours” into “inner-race bearing fault, developing since March”, ideally weeks before the stop.
- The fix closes the loop — someone root-causes the fault, engineers it out and verifies the repair in the data, so the bar on next quarter's Pareto is actually shorter.
That last step is the one no dashboard — uptime or condition — performs by itself, and it's where most monitoring investments quietly stall. A diagnosis that nobody actions is just a more detailed way of watching the machine fail. That's the gap AWI exists to close: software that reads the data you already own, and engineers who come to site to find the cause and engineer it out.
Where to Start
If you already track uptime, you're holding the map. Pull up the downtime Pareto, find the breakdown bar, and pick the machine that owns most of it — not your most critical asset, your most annoying one. Then go one layer deeper on that single machine: get its condition data read, its fault diagnosed, its cause engineered out, and the fix verified.
Here's the test worth remembering: an uptime pilot ends with a chart of your losses. A condition-monitoring pilot ends with a diagnosed machine. A closed-loop pilot ends with a fixed one. We record what that looks like end to end — signal detected, problem diagnosed, action taken, result verified. Only one of those three pilots changes next month's number — and that's the pilot we're ready to run, on your real plant data, now.
Frequently Asked Questions
What is the difference between uptime monitoring and condition monitoring?
Uptime monitoring measures whether a machine is running and for how long — it produces availability figures, OEE inputs and downtime totals. Condition monitoring measures the health of the machine itself, using signals like vibration, temperature, current and fluid analysis. Uptime monitoring tells you how much production you lost; condition monitoring tells you why the machine failed — and can flag the failure developing before any downtime exists to count.
Don't downtime reason codes already tell us why?
Only partly. Reason codes sort downtime into buckets — changeover, no material, breakdown — which is genuinely useful for the organisational causes. But for equipment failures, “breakdown” is a label, not a reason. The actual reason is physical: a bearing degrading, a misaligned coupling, contaminated oil. Reason codes are also typed in by busy operators, so they're opinion recorded after the event — where condition data is measurement, recorded before and during it.
Should we replace our machine monitoring or OEE system with condition monitoring?
No — they're layers of the same stack, not rivals. Uptime data sizes the loss and tells you which machines deserve attention; condition monitoring explains and predicts the breakdown share of that loss. The best sequence is to use the uptime data you already have to pick your worst asset, then add condition monitoring to that asset to find — and fix — the cause.
Can condition monitoring really catch failures before downtime happens?
Yes — that's its defining property. Most mechanical failures develop over weeks or months, and measurable symptoms such as rising vibration, temperature or wear-metal levels appear well before the machine stops. Uptime data can only record a failure after it has already cost you production; condition data is the leading indicator that lets you intervene while the machine is still running.
Key Takeaways
- Uptime data tells you the cost, never the cause. It's the scoreboard — essential for sizing the loss, unable to change it.
- The ladder is That → Where → Why → Never again — uptime, reason codes, condition monitoring, the fix. Each rung is only as useful as the one below it.
- “Breakdown” is a label, not a reason — reason codes are opinion after the event; condition data is measured physics from before it.
- Tense is the real difference — uptime data is past-tense (it counts downtime that already happened); condition data is future-tense (it warns while the machine still runs).
- It's a stack, not a rivalry — keep your uptime tool; use its Pareto to pick the asset, then go a layer deeper on that one machine.
- Only a closed loop changes next month's number — an uptime pilot ends with a chart; a condition pilot ends with a diagnosis; a closed-loop pilot ends with a fixed machine.
- Fluke Corporation / Censuswide (2025). Unplanned downtime costs UK manufacturers up to £736 million per week. digit.fyi — Fluke Corporation survey
- OEE.com (Vorne). OEE factors — availability, performance and quality, and how availability loss is measured. oee.com — OEE factors
- International Organization for Standardization. ISO 17359:2018 — "Condition monitoring and diagnostics of machines — General guidelines". iso.org — ISO 17359:2018
- International Organization for Standardization. ISO 15243:2017 — "Rolling bearings — Damage and failures — Terms, characteristics and causes" (failure-mode classification behind root-cause diagnosis). iso.org — ISO 15243:2017
- Deloitte. "Predictive Maintenance and the Smart Factory" — uptime gains of 10–20% and maintenance cost reductions of 5–10% from programmes that act on predictions. deloitte.com — Predictive Maintenance and the Smart Factory