Predictive maintenance process flowchart (condition-based)
Predictive maintenance flowchart template: monitoring setup and baseline, route or sensor data, alarm screening, diagnosis and confirmation, remaining-life estimate, planned work order and post-repair check.
What the predictive maintenance process flowchart (condition-based) process is
Predictive maintenance, also called condition-based maintenance, schedules work from evidence rather than from a date. The asset is instrumented or put on a monitoring route, each measurement is compared with the machine's own baseline, and the trigger for work is a change in condition: a vibration alarm, a hot joint on a thermal image, wear metals climbing in an oil sample, a bearing beginning to sing under ultrasound. The chart below follows one alarm end to end: the criticality review that decided this asset was worth monitoring at all, the limits set against its baseline, the reading collected on route or streamed from a sensor, the duty check that keeps the trend honest, the alarm screen, the diagnosis and its confirmation by a second technique, the remaining-life estimate, the condition report, the planned work order, the repair, the post-repair reading that proves the fault has gone, and the limits re-set from what was actually found inside the machine.
This is the condition-triggered case only. There is no calendar interval and no meter target at the top of the chart, so time-based and usage-based servicing belongs to preventive maintenance, which starts from the asset register rather than from a measurement. It is not the repair process for a defect somebody has already reported, and it is not the investigation that runs after a machine has stopped: this chart ends at a repair carried out before functional failure, and hands anything already broken to corrective maintenance and to failure investigation. The boundary matters because the entire economic argument for monitoring is that the intervention is planned; once the machine has failed, the reading that predicted it is history and the process is a different one. Treat the chart as a starting point to adapt under your own procedures, the manufacturer's limits and a competent analyst's review, particularly the branch that lets a machine keep running on a known fault, which is a safety decision as much as a maintenance one.
Four decisions carry the process. 'Machine on normal duty when read?' sits in the Operations lane because the person who knows the machine was running at half load is the operator, not the analyst, and a trend built from readings taken under different conditions raises alarms nobody can diagnose. 'Fault confirmed by a second technique?' is the false-alarm gate: one technique raises the question and a second answers it, and the unconfirmed branch loops back to the limit rather than forward to a work order. 'Can the asset run to a planned slot?' is the decision the whole programme exists to make, which is why it sits with the reliability engineer rather than with the planner who has to resource it. And 'Back to the baseline signature?' closes the loop at the end, because a repair that leaves the signature exactly where it was has not fixed the fault the report named.
What this flowchart covers
In this template
- Five swimlanes (Reliability engineer, Condition monitoring analyst, Operations, Maintenance planner and Technician) across six phases: set up monitoring, collect data, screen the alarm, diagnose, plan and repair, and verify and close
- The setup phase most condition monitoring charts leave out: a criticality review that decides which assets are monitored at all, a technique matched to each failure mode, a baseline taken on a healthy machine, and alert and danger limits set per measurement point
- A 'Route readings or continuous sensors?' decision at the top, so one chart covers both the periodic walk with a data collector and permanently installed sensors streaming to a platform, with the same screening and diagnosis behind either route
- A duty check in the Operations lane, 'Machine on normal duty when read?', whose off-duty branch sends the reading back to be collected again rather than letting a half-load measurement raise an alarm nobody can diagnose
- A 'Fault confirmed by a second technique?' gate that separates a real fault from a limit set too tight, where the unconfirmed branch closes the alarm and re-tunes the limit instead of raising work against the asset
- The planning half: a remaining-life estimate, a condition report, a 'Can the asset run to a planned slot?' decision with a cut-load-or-stop-now branch, a parts and window check that loops while long-lead parts are ordered, and a post-repair reading compared with the baseline before the catch is recorded
When to use this template
- You are standing up a condition monitoring programme and need the route, the analysis and the work order to be one process rather than three teams' spreadsheets
- Alarms are being raised and closed without work orders, and nobody can say how many of them turned out to be real faults
- You are moving critical assets off a fixed PM interval onto condition-based triggers and want the route agreed before any interval is cut
- Production and maintenance disagree about whether a machine with a known fault can run to the next shutdown, so the run-on decision needs a written rule and an owner
- You are fitting sensors or buying a monitoring platform and want the human decisions mapped before the data starts arriving
How it works
Rename the lanes to your roles
Replace Reliability engineer, Condition monitoring analyst, Operations, Maintenance planner and Technician with the roles you actually have. On many sites the analyst is a contractor who visits monthly and the reliability engineer is the maintenance manager wearing a second hat: merge those lanes rather than drawing a handoff that never happens.
List your techniques against your failure modes
Write on the chart which assets are on vibration, thermography, oil analysis, ultrasound or process data, and which failure mode each technique is there to catch. A technique chosen because the instrument was already owned, rather than because it detects the failure that actually occurs on that machine, produces readings nobody acts on.
Set the baseline and the limits per point
Record where the baseline came from, who may change an alert or danger limit, and what evidence justifies a change. Published severity bands are a starting point; the useful limit is the one set against this machine's own baseline at its normal duty, and it should move as the trend teaches you what normal looks like on that asset.
Fix the reading interval against the P-F window
State how often each point is read and why. The interval has to be short enough that at least one reading falls between the first detectable sign of a fault and the point at which the machine can no longer do its job, which is exactly why continuous sensors earn their place on assets whose warning period is short.
Write the confirmation and severity rules
Agree what counts as confirmation of a fault, which second technique is used for which failure mode, and who grades severity. Then decide what an unconfirmed alarm does: closing it silently loses the evidence, and closing it with a recorded limit change is what stops the same measurement point crying wolf every month.
Agree the run-on rule and who signs it
The 'Can the asset run to a planned slot?' branch needs a written rule and a named owner. Record which conditions force an immediate stop or a load reduction, what the remaining-life estimate must say before a machine is allowed to run to a shutdown, and how often that estimate is refreshed while it runs on.
Walk it against a real alarm
Take two or three alarms from the last year, one that became a work order and one that was closed as nothing, and trace them through the chart. Any step people describe that is not drawn, and any drawn step that nobody remembers doing, is the finding worth acting on before you publish it.
Frequently asked questions
What are the steps in a predictive maintenance process?
A criticality review picks the assets worth monitoring, a technique is matched to each failure mode, a baseline is taken on a healthy machine, and alert and danger limits are set per measurement point. Readings then arrive from a walked route or from installed sensors, and the operating state is checked before a reading is trusted. Within limits, it is logged to the trend and nothing else happens. Over the alert limit, the analyst works the spectrum and the fault pattern and confirms the fault with a second technique; an unconfirmed alarm is closed and the limit re-tuned. A confirmed fault is graded, its remaining life estimated and a condition report issued. The engineer decides whether the asset can run to a planned slot, the planner raises the work order and checks parts and window, the repair is done, a post-repair reading is compared with the baseline, and the catch is recorded.
What is the difference between predictive and preventive maintenance?
The difference is the trigger, not the planning. Preventive maintenance is due on a calendar interval or a meter target and is carried out whether or not anything is wrong, which makes it easy to schedule and means some of the work is done on machines that did not need it. Predictive or condition-based maintenance is triggered by a measurement showing that a fault has begun: vibration, temperature, oil condition, ultrasound or a process parameter drifting away from the machine's baseline. Both produce planned work orders and both are distinct from reactive maintenance, which starts after the machine has already stopped. Most sites run all three at once, with condition monitoring reserved for assets that are critical, that fail in a detectable way, and that give enough warning for anyone to act on.
What is the P-F interval and how does it set the monitoring frequency?
In reliability-centred maintenance the P point is the moment a developing fault first becomes detectable and the F point is functional failure, where the asset can no longer do what is asked of it. The time between them is the P-F interval, and it is the window in which a condition-based task can achieve anything. The practical rule is that readings must be more frequent than that window, so that at least one falls inside it; a common convention is to check at roughly half the P-F interval, which gives a margin for a missed reading. The window is a property of the failure mode and the technique, not of the asset, so the same bearing may give months of warning on oil analysis and days on vibration. Where the window is short, continuous monitoring is what makes the task worth doing at all. Treat any interval on this chart as a placeholder to replace with your own analysis.
Which condition monitoring techniques should we start with?
Start from the failure modes you actually see, not from the instrument catalogue. Vibration analysis covers bearing wear, imbalance, misalignment and looseness on rotating equipment, and machine vibration is evaluated in ISO 20816, whose zones run from newly commissioned through acceptable for long-term operation to levels needing immediate action. Infrared thermography, covered by ISO 18434, finds loose or corroded electrical connections and overheating parts. Oil analysis reports wear debris and contamination, coded for solid particles under ISO 4406. Ultrasound picks up leaks and early lubrication faults. ISO 17359 gives the general procedure for setting up a condition monitoring programme and its alarm criteria, and the ISO 18436 series sets requirements for qualifying the people who take and interpret readings. Read the standards themselves rather than a vendor summary of them.
How do you stop false alarms without losing the real ones?
Separate the three things that get called a false alarm. A data problem is a reading taken at the wrong speed or load, from a loose accelerometer or a drifting sensor, and it is caught by the duty check and by fixing the collection rather than the limit. A limit problem is a threshold set from a published band instead of from this machine's baseline, and it is fixed by re-tuning the limit with a named approver and a recorded reason. A genuine early detection is neither, and closing it as noise is how a programme loses its credibility. Confirming with a second technique is what tells them apart in the moment. Over a longer period, record every alarm and its outcome so you can report how many became work orders and how many were closed, and never delete the history: the alarms that turned out to be nothing are the evidence you need to justify moving a limit.