Incident severity classification flowchart

An incident severity classification flowchart: a decision tree taking availability, scope, business impact and exposure tests through to P1, P2, P3 or P4.

How it works

  1. Rename the lanes to your real decision rights

    Replace service desk, service owner, incident manager, and legal and compliance with the roles that actually hold each call in your organisation. If nobody owns the data or safety question out of hours, that is a rota gap to close before you redraw the lane.

  2. Write your scope thresholds onto the chart

    Open 'How many users or sites affected?' and replace the note with your own definitions of multiple sites, one team and single user: a region, a customer tier, a percentage of active users, a named account list. Thresholds that are not written down get renegotiated on every incident.

  3. Define unavailable versus degraded for your services

    Agree what 'Unavailable' means service by service, including partial availability cases such as read-only mode, a failed region, or a queue that is still accepting work but not processing it. Ambiguity here shifts the whole tree by one level.

  4. Set the workaround test honestly

    Decide what makes a workaround viable: documented, permitted, within capacity, and usable by the affected users today. A manual fallback that needs training, extra staff or a policy exception is not a workaround, and treating it as one is the most common way a P1 gets recorded as a P2.

  5. Attach a distinct response commitment to each outcome

    Each of 'P1 critical, bridge running', 'P2 high, fix under way', 'P3 medium, deadline tracked' and 'P4 low, scheduled work' should carry its own paging rule, update cadence and target times. If two levels produce identical behaviour, merge them rather than keeping a level nobody can distinguish.

  6. Agree who can reclassify, and record the reason

    The 'Impact changed at the next update?' decision is the only sanctioned route to change a level. Name who may take it, require the tests to be re-run rather than the level renegotiated, and record the new evidence and the time so the post-incident review can see when the picture changed.

  7. Circulate the chart for approval and keep the version

    Severity criteria are only useful if they are the agreed ones. Share the chart with service management, the service owners and legal for sign-off, then keep the approved version so the definitions you exercise are the definitions you publish.

Frequently asked questions

What is the difference between incident severity and incident priority?

Severity is a statement about impact: how much of the service is broken, for how many people, and what that blocks. Priority is a statement about ordering: what the team picks up next. ITIL does not actually use 'severity' as a formal term; it derives priority from impact and urgency. Many engineering and SRE teams use severity as shorthand for the impact half, which is how this chart uses it. The practical rule is that severity drives the response you owe (paging, bridge, update cadence) and only changes when the evidence changes, while priority drives the queue and can be re-set daily. Keeping one field for each stops people renegotiating the impact assessment in order to get faster service.

How is this different from an incident management process flowchart?

They answer different questions about the same subject. This chart is a decision tree: which severity level does this incident get, and who is entitled to decide. It is a chain of tests ending in named outcomes, and it stops as soon as the level is set. An incident management process flowchart is a cross-functional process map: what happens next, and who does it, from logging through diagnosis, escalation, resolution and closure. Most teams need both, and they connect at exactly one point, the step where a level is assigned. The process version is at /templates/incident-management-process.

How many severity levels should we have?

Four is the common default and the one this chart uses, with some organisations adding a P0 or SEV0 above it for existential events. The number matters less than whether each level carries a distinct, written response. If P3 and P4 lead to the same paging rule, the same target times and the same update cadence, you have three levels and an unused label. Fewer levels applied consistently beat more levels applied loosely, because the value of the classification is that everyone downstream can act on it without asking.

Who is allowed to declare a P1?

Name the role in advance and keep the list short. In this chart the declaration sits in the incident manager lane and three separate branches feed into it: a multi-site outage, a blocked critical function with no viable workaround, and a confirmed data or safety exposure. That structure is deliberate, because a P1 should be reachable by evidence from more than one direction, but declarable by a role that can also commit the response. Anyone should be able to request a P1; a named role confirms it.

Does a data or safety exposure automatically make an incident a P1?

In this chart, yes, and the exposure branch bypasses the user-count test entirely. The reason is that an exposure starts a clock that runs on awareness rather than on resolution, so a small incident can carry a large obligation. Under UK and EU GDPR, a personal data breach must be notified to the supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware of it, with notification to affected individuals a separate test tied to high risk. Other regimes and contracts set their own windows, and customer notice periods are often shorter than the statutory ones, so set the ones that apply to you on the decision box. Many organisations also route exposures out of this tree entirely into a security incident response process.

When should a severity level be changed after it has been set?

When the evidence changes, not when the pressure changes. This chart gives that a single route, the 'Impact changed at the next update?' decision, whose 'Impact grew' branch sends a P2 back into the P1 declaration. Re-run the tests rather than reopening the argument, and record what new evidence arrived and when. Downgrades follow the same rule and are worth handling explicitly, because an incident that stays at P1 after the impact has receded quietly trains people to ignore the level.

Use this template

More in IT process templates

More in Process map templates

Browse all IT process templates