There is a well-worn line in reliability marketing: catch it early and you save a fortune. It is true, but it is also where most of the industry stops helping.

An alert is a software output while a remediation project is a scope of work, a parts list with lead times, an isolation, a production window, a competent person to do the work, and a record that somebody will read in two years' time when the same asset misbehaves again. The distance between those two things is where reliability programmes quietly fail; not because the detection was wrong, but because either this is left to the manufacturer who lacks the time or expertise to complete the project, or the alerts provider is just a software company who cannot help with this bit.

This piece walks that distance in order: what a finding has to become before it can be scheduled, who holds direction and who holds the spanner, what stays with your team no matter who you hire, and what has to be in the record at the end. It is deliberately a piece about boundaries. The useful part is not the list of things we take on — it is the list of things we do not.

Circular diagram of six numbered stages — connect, monitor, detect, diagnose, act, verify — with a dashed branch leaving after detect, marking where a detect-and-alert tool stops and hands the fault back.
WHERE THE HANDOVER SITSThe six steps, with the branch where an alert is handed back to the maintenance team and the loop is left open. Everything in this article happens between diagnose and verify — the part of the ring most of the market does not sell.

A finding is not a job

A condition-monitoring finding reads something like this: the drive-end bearing on this pump is showing a defect signature that has been developing over several weeks, the trend is continuing, and the current-per-flow ratio has moved with it.

That is a good finding. It is not yet a job. Before anyone can act on it, a set of questions have to be answered, and none of them are measurements:

  • What does the failure mode actually call for — replacement, realignment, relubrication, remounting, or a change to how the asset is being run?
  • Can the work be done with the asset in place, or does it need to come out?
  • What parts are involved, and what are the lead times on them?
  • Who is competent to carry out the work, and are they in-house or contracted?
  • What does this asset do for production, and what does stopping it cost?
  • How long can it safely be left as it is, and what changes if it is left longer?

That last question is the one that turns a finding into a plan. The P-F curve is the usual way to think about it: there is an interval between the point a fault becomes detectable and the point it becomes a failure, and everything in this article is an argument about how to spend that interval well.

Detection buys you the interval. It does not spend it for you.

Not every finding becomes a project

It is worth saying plainly, because the opposite is what a supplier is incentivised to imply: most findings do not turn into remediation projects, and they should not.

A finding lands in one of four places, and deciding which is the first real piece of engineering judgement in the loop:

  • Watch it. The change is real but slow, the interval is long, and the right answer is to keep measuring and re-assess. This is the most common outcome and the least discussed.
  • Fold it into planned work. There is a shutdown in six weeks and the job belongs in it. No separate project, no separate cost.
  • Change how the asset is run. Sometimes the fault is a duty point, a start frequency or a process condition rather than a component. The cheapest remediation is the one where nothing is replaced.
  • Raise a project. The work needs scoping, parts, a window and direction, and it will not fit inside routine maintenance.

A supplier who takes every finding to the fourth option is not running a reliability programme, they are running a sales funnel. The point of condition data is to make the first three options available, because they are the ones that cost nothing.

From finding to scope of work

A scope of work is the document that makes a finding actionable. It is not a quotation and it is not a report. It states what is to be done, to which asset, under what conditions, by whom, and what evidence will be produced at the end.

Writing one properly needs three inputs. The first is the diagnosis — what is wrong and why, established at the machine rather than inferred from a dashboard. The second is the asset's history: what has already been done to it, how often the same job has come back, what was fitted last time and by whom. The third is the operational reality — the shift pattern, the planned stops, the spares already on the shelf, the access equipment available.

Most of that third category is knowledge your own team holds and nobody else does. A supplier who writes a scope of work without asking for it is writing fiction.

What a scope should pin down before anyone prices it:

  • The asset and its boundary. Named down to the individual unit, with the extent of work stated — the bearing, or the bearing and the shaft, or the whole drive end.
  • The intervention. What is being changed physically, and what is explicitly not in scope.
  • The condition for starting. Isolation, permits, access, and who issues them.
  • The window. When the asset is available, and how long it can be down.
  • The evidence. Which measurement will be re-taken afterwards, from which points, and against what baseline.

That last line is the one most scopes leave out, and it is the one that decides whether the job can be verified later. Agree it before the work starts, not after.

Who holds direction, and who holds the spanner

This is the part that is usually left vague, so we will be specific about our own arrangement.

AWI runs the engineering side of a remediation project: the diagnosis, the scope of work, the sequencing, the technical direction on site, and the verification at the end. Depending on the job and what you would rather keep in-house, the physical work is carried out either by your own fitters working to that scope, or by contractors we bring in and direct.

Both are legitimate. Which one applies is a commercial and practical decision, not a technical one, and it is worth settling early because it changes the scope, the price and the programme.

What does not change is where the boundaries sit. These stay with you, as the dutyholder and the operator of the plant, on every job:

  • Permits, isolation and safe systems of work. Your site, your procedures, your authorised persons.
  • Production decisions. When the asset can stop, and for how long, is an operations call.
  • Statutory inspection and examination. Condition monitoring is not a substitute for it and does not discharge it.
  • Acceptance. You decide whether the work is complete and whether the asset goes back into service.

And the things we do not take on at all, which is the more useful list:

  • We do not write to your control system. The monitoring path is read-only at plant control, and a remediation project does not change that.
  • We do not carry out statutory thorough examination.
  • We do not sell the sensors, the gateways or the parts. We are hardware-neutral, and that only means anything if we are not taking a margin on what we specify.
  • We do not run techniques we cannot stand behind in-house. Vibration, surface temperature and motor current are ours. Ultrasound, lubricant analysis and thermographic survey are brought in with specialists we work with, briefed on what the data already shows.

What actually stops a remediation project

In our experience the technical diagnosis is rarely the hard part. The things that stall a job are almost always logistical, and they are predictable enough to plan around:

  • Parts lead time. A bearing is days. A gearbox, a specified motor or anything with a long supply chain can be weeks or months. This is the single most common reason a correctly detected fault still runs to failure.
  • Access. Work at height, confined space, or a machine that cannot be reached without moving something else.
  • The window. The asset is available in three weeks and the parts arrive in five.
  • Competence and cover. The person who knows this machine is on annual leave the week of the planned stop.
  • The asset getting worse while you wait. Which is why monitoring continues through the waiting period rather than stopping at the alert.

None of that is glamorous. All of it is the difference between a programme that closes jobs and one that accumulates findings nobody has actioned — which is its own failure mode, and one that quietly teaches a maintenance team to stop reading the alerts.

Closing it out: what has to be in the record

A remediation project is not finished when the work is done. It is finished when the record is complete, and the record has to be built as the job runs rather than written up afterwards from memory.

We keep every job to the same four fields, and each one defends against a specific way that maintenance records go wrong:

Signal detected

The asset, the measurement, and what moved — captured as it was when it triggered, not reconstructed later. This defends against hindsight: it is very easy, once you know what the fault was, to remember the data as clearer than it was.

Problem diagnosed

What was found on the machine, and the evidence behind the conclusion. This defends against the symptom being recorded as the cause. If the alert described one thing and the strip-down found another, the record says so — that disagreement is one of the most useful things in the file.

Action taken

What changed physically, who did it, and when. This defends against the most common gap in a CMMS: a work order closed with a line of text that tells the next engineer nothing.

Result verified

The same measurement points, read again after the work, against the agreed baseline. This defends against the assumption that a completed job is a successful one. It is also the stage most reporting leaves out entirely.

Time-series diagram of one measurement: a steady baseline band, a rising trend crossing a configured limit, a marked intervention window, then a return to baseline read at the same points.
WHAT CLOSING OUT LOOKS LIKEThe same sensor and the same measure, read before the intervention and again in the verification window afterwards. The comparison is only possible if the measurement points and the baseline were agreed in the scope of work.

A note on how we publish these. Our intervention record template is deliberately empty. We have not filled it with an illustrative job, because a worked example with a machine name and a set of figures reads as a customer whether or not it is labelled as one. When a pilot loop closes and the write-up is agreed with the customer, it will be published with their name on it or not at all.

Why this is the half that gets skipped

Detection is a software problem and it is getting cheaper every year. Scoping, sequencing, directing and verifying physical work on somebody else's plant is a people problem, and it does not scale the way software does. That is why the market has concentrated on the first half and handed the second back to you.

If you already have a strong maintenance team, an engineering manager with capacity, and spares discipline, handing the second half back may be exactly what you want — and a monitoring platform on its own is a reasonable purchase. Plenty of SME manufacturers do not have all three, and for them a detection tool alone produces a queue of findings and no change in the failure rate.

That is the whole argument for closing the loop rather than stopping at the alert. Not that detection is unimportant — it is the thing that buys the interval — but that an interval you do not spend is worth nothing.

Key takeaways

  • An alert is a software output; a remediation project is a scope of work, a parts lead time, an isolation, a window and a record. The gap between them is where programmes fail.
  • A scope of work should name the asset boundary, the intervention, the conditions for starting, the window, and the evidence that will be produced — agreed before the job, not after.
  • Permits, isolation, production decisions, statutory inspection and final acceptance stay with the dutyholder whoever carries out the work.
  • Parts lead time is the most common reason a correctly detected fault still runs to failure. Monitoring should continue through the wait.
  • The record is the deliverable. Four fields, built as the job runs: signal detected, problem diagnosed, action taken, result verified.
  • Detection buys you an interval. Spending it is a different discipline, and it is the one that changes the failure rate.
  1. International Organization for Standardization. ISO 17359:2018 — "Condition monitoring and diagnostics of machines — General guidelines". iso.org — ISO 17359:2018
  2. International Organization for Standardization. ISO 13379-1:2012 — "Condition monitoring and diagnostics of machines — Data interpretation and diagnostics techniques — Part 1: General guidelines". iso.org — ISO 13379-1:2012