Equipment breakdown process flowchart (unplanned failure)
Equipment breakdown process flowchart for unplanned machine failure: safe stoppage, downtime clock, repair-on-site and spares decisions, test and restart.
What the equipment breakdown process flowchart (unplanned failure) process is
An equipment breakdown is an unplanned failure that stops a machine and usually stops production with it. The process below is the reactive route back: make the machine safe, get the fault reported and the downtime clock running, assess it, decide whether it can be fixed on site, find the part, repair, prove the machine still runs to specification, and hand it back. Every step is on the critical path, because until the last one finishes the line is not making anything.
It is worth being clear about what this process is not. It is not the general maintenance work-order process: planned and condition-based work is scheduled into a window, goes through job planning and permit issue, and closes out with labour and cost capture, none of which applies when a machine has already stopped. It is not IT incident management, which restores a service rather than a machine and has no isolation or spares step. And it is not the investigation itself. This chart ends at production resumed and refers repeat failures out to root cause analysis and corrective action, rather than holding the line down while an engineer works out why the bearing went. Maintenance terminology standards such as EN 13306 draw the same line, treating work carried out after a fault has been recognised as corrective maintenance and keeping it separate from preventive work.
The template maps the flow across five lanes (Operator, Shift supervisor, Maintenance technician, Stores and Engineering) and five phases from the stoppage to restart and review. It keeps the branches most breakdown procedures leave undocumented: what happens when the job cannot be done on site, what happens when the spare is not on the shelf, whether a temporary fix is allowed and how it gets followed up, and what triggers a root cause referral rather than another repair.
What this flowchart covers
In this template
- Five swimlanes with a named owner for every step (Operator, Shift supervisor, Maintenance technician, Stores and Engineering) across five phases: Stoppage and report, Assessment, Parts, Repair, and Restart and review.
- The safe stoppage sequence in the Operator lane: 'Breakdown detected on the line' runs into 'Stop machine and make it safe' before 'Report fault to shift supervisor', so isolation happens ahead of the paperwork rather than alongside it.
- Downtime capture at the handover point ('Log fault and start downtime clock' followed by 'Assign technician and raise work order' in the Shift supervisor lane) so the clock starts at the stoppage rather than when a technician arrives.
- A 'Repairable on site?' decision whose no branch goes to Engineering for 'Arrange contractor or unit replacement' and rejoins the same functional test, so external work is not an undocumented side path.
- A 'Spare part in stock?' decision in the Stores lane with both outcomes drawn: issue the part and update the stock record, or order it with a confirmed lead time and then take the 'Temporary fix possible?' decision, which either applies a temporary fix with a flagged follow-up or waits for the part.
- Verification and close-out: 'Run functional test on the machine' feeds an 'Operating to specification?' decision that loops back to assessment when it fails, then 'Restart production and record downtime' and a 'Repeat failure on this asset?' decision that refers the asset to Engineering for root cause before 'Production resumed'.
When to use this template
- Writing or refreshing the breakdown procedure for a plant or line, so everyone can see who stops the machine, who logs it and who authorises the restart.
- Downtime figures are not trusted, because the clock starts and stops at different points depending on who filled in the sheet.
- Configuring a CMMS or maintenance module: the work order, failure code, parts issue and downtime fields in the chart map onto the fields you have to set up.
- Temporary fixes are quietly becoming permanent, and you want the follow-up work order to be part of the process rather than someone's memory.
- Inducting operators and new technicians who need to know what they may do themselves and exactly where the handover to maintenance sits.
How it works
Rename the lanes to your real roles
Replace Operator, Shift supervisor, Maintenance technician, Stores and Engineering with the roles you actually have. Smaller sites merge Engineering into the technician lane; sites running around the clock often replace Shift supervisor with the shift lead or the control room. Merge a lane rather than leaving one that only a single step ever touches.
Attach your safe isolation rules to the stop step
'Stop machine and make it safe' is where your safe system of work belongs: emergency stop, isolation and lock-off, any permit needed, and who is competent to do it. In Great Britain the Provision and Use of Work Equipment Regulations 1998 require work equipment to be capable of being isolated from its sources of energy. The diagram should point at your existing isolation procedure rather than restate it, so the two cannot drift apart.
Define when the downtime clock starts and stops
Write down the two timestamps: the clock starts when the machine stopped and ends when it is producing acceptable output again. Most arguments about downtime data come from the start point, so fix it once and use the same definition on every asset. Decide at the same time whether you record waiting-for-spares time separately, since that is usually the largest and most fixable component.
Set the on-site repair boundary
Say what makes 'Repairable on site?' a no: specialist tooling, a warranty seal, a sealed or calibrated unit, statutory inspection, or work nobody on shift is competent to carry out. Then name who can authorise a contractor call-out or a replacement unit, and up to what value, so the escalation does not stall waiting for someone to be found.
Write the temporary fix rules down
Decide who may authorise a temporary fix, what may never be worked around under any circumstances — guarding, interlocks and other safety functions — what output or product restriction applies while it is in place, and how the permanent repair gets its own dated work order. Without those four answers the branch becomes an informal habit instead of a controlled decision.
Set the repeat failure trigger, then publish a versioned copy
Define what makes 'Repeat failure on this asset?' a yes, such as the same failure mode on the same asset inside an agreed period, and name where the referral goes. Walk the finished chart with each shift and correct the steps they actually perform, then publish it as the current version with a sign-off so everyone is working from the same revision.
Frequently asked questions
What are the steps in an equipment breakdown process?
Detect the failure and stop the machine safely, report the fault, log it with the downtime start and raise a work order, assess the fault, decide whether it can be repaired on site, obtain the spare or arrange external work, carry out the repair, run a functional test, confirm the machine is operating to specification, restart production and record the downtime, then refer repeat failures for root cause analysis. The order matters in two places in particular: making the machine safe comes before reporting it, and the functional test comes before handback rather than after production has already restarted.
What is the difference between breakdown maintenance and preventive maintenance?
Breakdown maintenance is carried out after a failure, when the machine has already stopped. Preventive maintenance is carried out to a schedule or to a measured condition before it fails. Maintenance terminology standards such as EN 13306 make the same split, treating work carried out after a fault has been recognised as corrective maintenance. The practical consequence is that breakdown work sits on the critical path, so this process is built around speed, escalation and a downtime clock, while planned work is scheduled into an agreed window and needs job planning and resource booking instead.
When should the downtime clock start and stop?
It should start when the machine stopped and stop when it is producing acceptable output again — not when the work order was raised, and not when the technician put the tools down. Anything narrower understates the loss, because waiting for a technician, waiting for a part and running the functional test are all time the line was not producing. Use one definition across every asset, since comparisons between lines are meaningless if one supervisor starts the clock at the stoppage and another at the call-out.
Is a temporary fix acceptable during a breakdown?
Sometimes, and the process should say so explicitly rather than leaving it to whoever is on shift. A temporary fix is a decision to run at reduced output, or with something worked around, until the correct part arrives. It should be authorised by the supervisor rather than agreed informally, must never involve bypassing guarding, interlocks or any other safety function, should carry a stated restriction on output or product, and should raise its own dated work order for the permanent repair. The risk this branch guards against is not the fix itself but the fix that is still in place a year later because nobody recorded it.
When should a breakdown be referred for root cause analysis?
When repairing it again is no longer the right answer: the same failure mode recurring on the same asset inside an agreed period, a failure with a safety or quality consequence, or one whose downtime crosses a threshold you have set. This process deliberately ends at production resumed and hands the investigation to engineering, so the line is not held down while somebody works out why the component failed. Referring the asset does not slow the restart; it makes sure the failure is counted and investigated rather than quietly absorbed into the shift.