Incident severity classification flowchart
An incident severity classification flowchart: a decision tree taking availability, scope, business impact and exposure tests through to P1, P2, P3 or P4.
What the incident severity classification process is
Incident severity classification is the step that turns a report into a level. Someone reads the evidence, applies an agreed set of tests, and lands on P1, P2, P3 or P4, and everything downstream keys off that answer. It is worth separating from priority at the outset, because the two words are used interchangeably and then argued about mid-incident. Severity describes impact and therefore the response you owe: who is paged, whether a bridge opens, how often stakeholders are updated. Priority describes ordering: what the team picks up next. Severity is assessed from evidence and only moves when the evidence moves; priority can be re-set every morning without anyone reopening the impact assessment.
This page is a decision tree, not a process map, and that is the reason it exists separately. It answers one question, which level does this incident get and who is entitled to decide, by chaining seven tests together until every path reaches a named outcome. It deliberately stops at the moment the level is set. If what you need is the end-to-end flow after that point, how the ticket is diagnosed, escalated, resolved, confirmed with the user and closed across the service desk, incident manager and support tiers, use the incident management process flowchart at /templates/incident-management-process. Use this chart to decide the level; use that one to run the work.
The swimlanes here name decision rights rather than departments, which is the part most severity matrices leave implicit. The service desk answers what can be observed from the report: is the service unavailable or degraded, and how many users or sites are affected. The service owner answers whether a critical business function or revenue path is genuinely blocked. The incident manager owns the workaround test, the declaration itself and the review. Legal and compliance own two questions and nothing else: is there a data or safety exposure, and is a regulatory or contractual clock running. Four bands is enough to settle the who-decides argument without drawing an organisation chart.
What this flowchart covers
In this template
- Four decision-rights lanes rather than a department map (service desk, service owner, incident manager, legal and compliance) laid out across five assessment stages: intake, scope of impact, business impact, severity call, and outcome and review.
- The opening test, 'Service unavailable or degraded?', with three answers: 'Unavailable' and 'Degraded' continue into the tree, while 'No impact' terminates straight away at 'Closed as a service request', so a request or query never picks up a severity level.
- Scope assessed before business impact: 'How many users or sites affected?' splits into 'Multiple sites', which goes directly to a P1 declaration, 'One team', which is passed to the service owner, and 'Single user', which is still checked for exposure.
- Two independent escalators to critical: 'Critical function or revenue blocked?' feeding 'Viable workaround available?', where 'No workaround' declares P1 and 'Workaround' assigns P2; and 'Data or safety exposure?', where 'Exposure' declares P1 however few users are affected.
- The low end of the matrix decided by one question, 'SLA or regulatory clock engaged?', separating 'P3 medium, deadline tracked' from 'P4 low, scheduled work' rather than leaving it to judgement, giving five distinct endpoints in all alongside 'P1 critical, bridge running', 'P2 high, fix under way' and the closed-as-a-request exit, plus an explicit reclassify path where 'Impact changed at the next update?' sends 'Impact grew' back into the P1 declaration.
- Written criteria on the pivotal decisions: what unavailable means, where the user and site thresholds sit, what counts as a viable workaround, and why an exposure is treated as critical regardless of scope.
When to use this template
- Writing down the severity definitions your team currently argues about, so two people triaging the same incident at three in the morning reach the same level without negotiating.
- Configuring an ITSM or on-call tool where severity is a required field, and you need the tests behind each level rather than a dropdown with four options and no guidance.
- Agreeing decision rights before an outage rather than during one: who may declare a P1, who confirms a data or safety exposure, and who is allowed to reclassify.
- Reviewing a past incident to check whether the level assigned matched the evidence available at the time, which is a different question from whether the response was good.
- Onboarding on-call engineers and new service desk agents who need one page showing the questions, the thresholds and where each answer lands.
How it works
Rename the lanes to your real decision rights
Replace service desk, service owner, incident manager, and legal and compliance with the roles that actually hold each call in your organisation. If nobody owns the data or safety question out of hours, that is a rota gap to close before you redraw the lane.
Write your scope thresholds onto the chart
Open 'How many users or sites affected?' and replace the note with your own definitions of multiple sites, one team and single user: a region, a customer tier, a percentage of active users, a named account list. Thresholds that are not written down get renegotiated on every incident.
Define unavailable versus degraded for your services
Agree what 'Unavailable' means service by service, including partial availability cases such as read-only mode, a failed region, or a queue that is still accepting work but not processing it. Ambiguity here shifts the whole tree by one level.
Set the workaround test honestly
Decide what makes a workaround viable: documented, permitted, within capacity, and usable by the affected users today. A manual fallback that needs training, extra staff or a policy exception is not a workaround, and treating it as one is the most common way a P1 gets recorded as a P2.
Attach a distinct response commitment to each outcome
Each of 'P1 critical, bridge running', 'P2 high, fix under way', 'P3 medium, deadline tracked' and 'P4 low, scheduled work' should carry its own paging rule, update cadence and target times. If two levels produce identical behaviour, merge them rather than keeping a level nobody can distinguish.
Agree who can reclassify, and record the reason
The 'Impact changed at the next update?' decision is the only sanctioned route to change a level. Name who may take it, require the tests to be re-run rather than the level renegotiated, and record the new evidence and the time so the post-incident review can see when the picture changed.
Circulate the chart for approval and keep the version
Severity criteria are only useful if they are the agreed ones. Share the chart with service management, the service owners and legal for sign-off, then keep the approved version so the definitions you exercise are the definitions you publish.
Frequently asked questions
What is the difference between incident severity and incident priority?
Severity is a statement about impact: how much of the service is broken, for how many people, and what that blocks. Priority is a statement about ordering: what the team picks up next. ITIL does not actually use 'severity' as a formal term; it derives priority from impact and urgency. Many engineering and SRE teams use severity as shorthand for the impact half, which is how this chart uses it. The practical rule is that severity drives the response you owe (paging, bridge, update cadence) and only changes when the evidence changes, while priority drives the queue and can be re-set daily. Keeping one field for each stops people renegotiating the impact assessment in order to get faster service.
How is this different from an incident management process flowchart?
They answer different questions about the same subject. This chart is a decision tree: which severity level does this incident get, and who is entitled to decide. It is a chain of tests ending in named outcomes, and it stops as soon as the level is set. An incident management process flowchart is a cross-functional process map: what happens next, and who does it, from logging through diagnosis, escalation, resolution and closure. Most teams need both, and they connect at exactly one point, the step where a level is assigned. The process version is at /templates/incident-management-process.
How many severity levels should we have?
Four is the common default and the one this chart uses, with some organisations adding a P0 or SEV0 above it for existential events. The number matters less than whether each level carries a distinct, written response. If P3 and P4 lead to the same paging rule, the same target times and the same update cadence, you have three levels and an unused label. Fewer levels applied consistently beat more levels applied loosely, because the value of the classification is that everyone downstream can act on it without asking.
Who is allowed to declare a P1?
Name the role in advance and keep the list short. In this chart the declaration sits in the incident manager lane and three separate branches feed into it: a multi-site outage, a blocked critical function with no viable workaround, and a confirmed data or safety exposure. That structure is deliberate, because a P1 should be reachable by evidence from more than one direction, but declarable by a role that can also commit the response. Anyone should be able to request a P1; a named role confirms it.
Does a data or safety exposure automatically make an incident a P1?
In this chart, yes, and the exposure branch bypasses the user-count test entirely. The reason is that an exposure starts a clock that runs on awareness rather than on resolution, so a small incident can carry a large obligation. Under UK and EU GDPR, a personal data breach must be notified to the supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware of it, with notification to affected individuals a separate test tied to high risk. Other regimes and contracts set their own windows, and customer notice periods are often shorter than the statutory ones, so set the ones that apply to you on the decision box. Many organisations also route exposures out of this tree entirely into a security incident response process.
When should a severity level be changed after it has been set?
When the evidence changes, not when the pressure changes. This chart gives that a single route, the 'Impact changed at the next update?' decision, whose 'Impact grew' branch sends a P2 back into the P1 declaration. Re-run the tests rather than reopening the argument, and record what new evidence arrived and when. Downgrades follow the same rule and are worth handling explicitly, because an incident that stays at P1 after the impact has receded quietly trains people to ignore the level.