AWI Ltd
Home AWI Analytics AWI Labs UK SMEs Calculator About Blog Contact
in Contact us
POSITIONING 20 JULY 2026 · 11 MIN READ

Closed-Loop Reliability: Why Detecting the Fault Is Only Half the Job

The reliability market has quietly solved the wrong half of the problem. Sensors are cheap, AI is capable, and almost any tool can now tell you a machine is about to fail. Then it sends an alert and hands the problem back. Detection has become the easy part — and the hard, valuable part is everything that happens after. This is the case for closing the loop.

This is the argument that sits underneath everything else we write — the vibration and fluid analysis explainers, the predictive maintenance guide, all of it. Every one of those techniques detects a problem. The question this article answers is: and then what?

Detection Used to Be the Hard Part. Now It's the Cheap Part.

Twenty years ago, knowing a bearing was degrading required a certified analyst with a handheld collector walking the plant on a route. Detection was genuinely difficult, genuinely scarce, and genuinely where the value sat. That world is gone.

Wireless vibration sensors have fallen from thousands of pounds to a few hundred per point. Machine-learning models learn an asset's normal signature automatically. Cloud platforms ingest historian data by the million rows. The result: detection has been commoditised. Monitoring software, digital twins, corrosion sensors, production-intelligence platforms — the whole category can now tell you, with reasonable accuracy, that something is wrong. The alert is table stakes.

And yet the headline problem hasn't moved. Unplanned downtime still costs UK manufacturers up to £736 million every week, and as of 2025 the majority still hadn't invested in predictive maintenance at all. [1] If detection were the bottleneck, cheaper detection would have fixed it. It hasn't — because the bottleneck was never detection.

< 1%
of the industrial data already collected is ever used to make a decision — McKinsey Global Institute [2]

Read that number the right way and it's not a data-collection problem — it's an action problem. The data exists. The detection works. What's missing is everything between the alert and the fix.

The Reliability Loop — and Where It Breaks

Preventing a failure from recurring isn't a single event. It's a loop of four stages, and it only pays back when it completes:

  • Detect — the data flags that something is drifting. The commoditised part.
  • Diagnose — work out the root cause, not just the symptom. Is the bearing failing, or is a misaligned coupling killing it?
  • Engineer it out — physically fix the cause so the same fault doesn't return: realign, reseal, refilter, rebuild, protect.
  • Verify — confirm in the data that the intervention worked, and feed that back as the new baseline. The loop closes.

International standards already frame monitoring this way — ISO 17359, the umbrella standard for condition monitoring, describes it as a full procedure from detection through to corrective action and review, not a standalone alarm. [3] The tools, almost universally, implement only the first step. How much warning that first step can give — and why it varies by failure mode — is the subject of the P-F curve.

WHERE MOST TOOLS STOP
Detect Alert — handed back to you
THE LOOP, CLOSED
Detect Diagnose Engineer it out Verify ↻ repeat

Why Software-Only Tools Stop at the Alert

This isn't a criticism of any one product — it's structural, and it's worth understanding, because it tells you what a category can't do no matter how good its models get.

Pure-software and sensor companies are built to be asset-light. Their entire economic advantage is that one platform serves ten thousand customers at near-zero marginal cost — no vans, no engineers, no travel, no physical work. That model is what makes them scalable and fundable, and it is precisely what stops them at the alert. The moment a problem needs a human standing on your shop floor with a spanner, it breaks their unit economics. So they don't do it — by design, not by oversight.

Which is fine, if you have a reliability team of your own to catch the handoff. Most manufacturers don't — and that's where the open loop turns from an inconvenience into the whole problem.

A detection tool sold to a team that has no one to action its alerts isn't a solution. It's a more sophisticated way of documenting failures you were always going to have.

What Closing the Loop Actually Takes

The three stages after detection each demand something a dashboard can't provide on its own.

Diagnose the cause, not the symptom

An alert says "bearing vibration rising." That's a symptom. The cause might be the bearing — or a misaligned coupling, a bent shaft, cavitation, contamination, or a soft foot in the baseplate. Swap the bearing and the real cause quietly kills the next one. Proper diagnosis reads the failure mode — the ISO 15243 classification of what's actually happening and why [4] — by correlating signals and, often, by someone physically inspecting the machine. This is where the data and the engineer meet.

Engineer it out

Diagnosis without action is just a better-informed breakdown. Closing the loop means physically removing the cause: realigning the coupling, correcting the lubrication regime, upgrading a seal to the right material, rebuilding and protecting an eroded surface rather than replacing the whole asset. It's hands-on engineering, on site, and it's the step the entire software market outsources back to you. That physical half of the loop is what AWI Labs exists to do — specifying the sensors, doing the diagnosis on site, and carrying out the repair.

Verify in the data

The loop only truly closes when the data confirms the fix worked — vibration back to baseline, temperature settled, the fault frequency gone. That verification becomes the new normal the system watches against, and the cycle starts again a little smarter. Without it, you're guessing whether the repair held; with it, reliability compounds. Every stage gets written down the same way — signal detected, problem diagnosed, action taken, result verified. That record is the evidence trail a reliability programme is judged on.

The Real Cost of an Open Loop

An open loop doesn't just fail to help — it actively decays:

  • Alert fatigue. Alerts that nobody can action become noise. Within months the team learns to ignore them, and the expensive detection layer goes quiet in everyone's inbox.
  • Repeat failures. Fix the symptom, not the cause, and the same fault returns — on the same asset, often on the same schedule — each time booked as a fresh "unplanned" event.
  • Shelfware. A monitoring platform that never changed a maintenance decision is the most expensive dashboard a factory owns. It's a large share of why so many digital projects quietly stall.
  • The downtime still happens. Detection that isn't actioned doesn't move the £736m-a-week number one penny. Only the completed loop does.

Deloitte's research on predictive maintenance puts the prize at 10–20% higher uptime and 5–10% lower maintenance cost — but those returns come from programmes that act on predictions, not ones that merely generate them. [5] The gap between the two is the open loop.

Closed-Loop Reliability Is an SME Problem First

A large plant with a dedicated reliability engineering department can, in principle, catch the handoff itself: the software flags, the in-house team diagnoses and fixes. The open loop is survivable because they own the missing stages.

An SME manufacturer has no such team. A senior data scientist costs £70k+; a full reliability function is out of reach entirely. So for an SME, the alert is the dead end — there's simply no one on staff whose job is to turn it into a fix. This is why bolting a detection tool onto an SME so often changes nothing, and why the SME case is different from the enterprise one. The thing an SME actually needs isn't more detection. It's the whole loop, delivered as a service.

Is Your Reliability Programme Open or Closed?

A quick self-diagnosis. Run your current setup down the two columns:

StageOpen loopClosed loop
DetectAlert firesAlert fires
DiagnoseLeft to the customerRoot cause found, on site
FixYour problemEngineered out by the provider
VerifyRarely happensConfirmed in the data
OutcomeThe same failure, againThe failure, gone for good

If more than one row in your operation lives in the left column, you have a detection layer, not a reliability programme — and it's one of the clearest signs your tools are holding you back.

Where to Start: Close the Loop Once

You don't prove closed-loop reliability with a platform tour. You prove it by closing the loop, once, on a real asset — the pump, motor or gearbox that keeps biting you. Read its existing data, diagnose the true cause, engineer it out, and verify the fix held. One completed cycle tells you more than any dashboard demo, because it's the whole value proposition compressed into a single asset.

That's exactly what AWI is built to do — software that reads the data you already own, and engineers who turn up and act on what it finds. If you want to size the prize first, the downtime cost calculator puts a number on what one closed loop is worth on your site.

Frequently Asked Questions

What is closed-loop reliability?

Closed-loop reliability is a maintenance approach that completes the full cycle after a fault is detected: detect the problem in the data, diagnose the true root cause, physically engineer that cause out on site, and verify in the data that the fix worked. Most monitoring tools only perform the first step — detection — and hand the remaining stages back to the customer, leaving the loop open.

Why isn't condition monitoring software enough on its own?

Condition monitoring software detects and alerts, but detection has become the cheap, commoditised part of reliability. The value is in diagnosing the root cause and physically fixing it — work that requires engineers on site. Software-only tools are built to be asset-light and remote, so they structurally can't perform those steps. For a team without its own reliability engineers, an unactioned alert changes nothing.

What does closing the loop require that software can't do alone?

Three things: root-cause diagnosis (distinguishing the failing component from the underlying cause, often needing physical inspection), hands-on engineering to remove that cause (realignment, resealing, rebuilding, protecting), and verification in the data that the intervention held. All three need people on the shop floor, not just a model in the cloud.

Is closed-loop reliability only for large manufacturers?

The opposite — it matters most to SMEs. Large plants often have in-house reliability teams to action alerts, so an open loop is survivable. SMEs rarely have that capacity, so for them the alert is a dead end. Closed-loop reliability delivered as a service — software plus engineers on site — is what makes predictive maintenance actually work for a smaller manufacturer.

Key Takeaways

  • Detection is commoditised. Cheap sensors and capable AI mean almost any tool can flag a fault. The alert is table stakes, not a differentiator.
  • The loop has four stages — detect, diagnose, engineer out, verify — and only pays back when it completes. Most tools stop at stage one.
  • Software-only stops at the alert by design — asset-light economics can't put an engineer on your floor, no matter how good the model.
  • An open loop decays into alert fatigue, repeat failures and shelfware — and never moves the downtime number.
  • It's an SME problem first. Without an in-house reliability team, the alert is the dead end — so SMEs need the whole loop delivered as a service.
  • Prove it by closing the loop once — on one real asset, end to end. That's worth more than any platform demo.
SOURCES & REFERENCES
Sources & References
  1. Fluke Corporation / Censuswide (2025). Unplanned downtime costs UK manufacturers up to £736 million per week; majority not yet invested in predictive maintenance. digit.fyi — Fluke Corporation survey
  2. McKinsey Global Institute. "The Internet of Things: Mapping the Value Beyond the Hype" — less than 1% of industrial data collected is used for decision-making. mckinsey.com — Internet of Things value
  3. International Organization for Standardization. ISO 17359:2018 — "Condition monitoring and diagnostics of machines — General guidelines" (monitoring as a full procedure through to corrective action). iso.org — ISO 17359:2018
  4. International Organization for Standardization. ISO 15243:2017 — "Rolling bearings — Damage and failures — Terms, characteristics and causes" (failure-mode classification for root-cause diagnosis). iso.org — ISO 15243:2017
  5. Deloitte. "Predictive Maintenance and the Smart Factory" — uptime gains of 10–20% and maintenance cost reductions of 5–10% from programmes that act on predictions. deloitte.com — Predictive Maintenance and the Smart Factory

Stop collecting alerts. Start closing the loop.

Detection, diagnosis, the physical work and the proof it held sit with one partner rather than four. AWI Analytics reads the plant data you already hold and says what is changing and why; AWI Labs engineers go to the machine, engineer the cause out and re-measure afterwards. AWI Analytics is open to a limited number of pilot partners, with a free 90-day pilot on your own plant data.
Contact us
← ALL ARTICLES
AWI Ltd
Home AWI Analytics AWI Labs UK SMEs Calculator About Blog Contact Privacy Security
in
© 2026 AWI LTD. ALL RIGHTS RESERVED. MAINTENANCE FROM REACTIVE TO PROACTIVE · DOWNTIME FROM RANDOM TO PREDICTABLE