Disaster recovery process flowchart (IT systems)

A swimlane flowchart of the IT disaster recovery process: invocation criteria, RTO and RPO prioritisation, failover, data restore, validation and failback.

How it works

  1. Rename the lanes to the roles you actually have

    Replace Detection / Operations, DR coordinator, Management, Infrastructure team and Application owners with your real roles: NOC or monitoring, IT service continuity manager, the executive on the crisis rota, platform and database teams, named service owners. Smaller organisations often merge the coordinator and infrastructure lanes; if a lane has nobody in it, delete it rather than leave it unstaffed.

  2. Write your invocation criteria onto the decision

    Make 'DR invocation criteria met?' objective enough for someone to apply under pressure: primary site inaccessible, expected outage longer than the RTO of a tier-one system, primary storage or database unrecoverable in place. Name the role authorised to invoke and a deputy, and record how they are reached outside working hours.

  3. Attach real RTO and RPO figures to each system

    Take the targets from the business impact analysis, not from what the infrastructure currently achieves, and list them in tiers against 'Prioritise systems by RTO and RPO'. Record the dependency order as well as the priority order, because a tier-one application cannot be validated before the identity, network and database services beneath it are up.

  4. Describe your actual recovery mechanism

    'Fail over to the recovery site' means something different for replicated infrastructure, a warm standby, a second cloud region and a rebuild from backup media. Put the mechanism, the runbook link, the location of credentials and break-glass accounts, and any manual DNS or network changes into the box comments, so the chart is usable during a recovery and not only during a review.

  5. Define what 'integrity verified' means and who says so

    Set the checks behind 'Data integrity verified?': record counts, consistency and referential checks, application-level smoke tests, comparison against the last known good state. Decide what 'an earlier copy' means, how many restore attempts you make before escalating, and note that each step back increases the data loss you will have to declare against the RPO.

  6. Add failback, then test the chart and keep the version

    Confirm the failback path matches how you work: window, data resynchronisation from the recovery site, change control approval, and the point at which DR is formally stood down. Then exercise the chart, compare the recovery time and data loss you achieved against your targets in the post-event review, and publish the approved version so people rehearse the same revision they would follow live.

Frequently asked questions

What is the difference between disaster recovery and business continuity?

Business continuity is about keeping the organisation delivering its products and services during a disruption, by any means available: manual workarounds, alternative premises, redeployed staff, standby suppliers, customer communications. Disaster recovery is the IT part of that: restoring the systems, data and infrastructure the organisation depends on. Continuity work sets the priorities, because the business impact analysis decides which activities matter most and how long they can be interrupted; disaster recovery inherits those timings as RTO and RPO targets and delivers against them. A DR plan without continuity input tends to recover whatever is easiest to recover first.

What do RTO and RPO actually mean?

RTO, the recovery time objective, is the target time within which a system or service must be usable again, measured forward from the disruption rather than from the moment somebody signs the invocation form. RPO, the recovery point objective, is the point in time that data must be restored to, measured backwards from the disruption, so in practice it is the maximum data loss the organisation is willing to accept. They drive different investments: RTO is set by how quickly you can bring infrastructure up (standby capacity, automation, rehearsal), while RPO is bounded by how often data is copied off, so nightly backups cannot support an RPO shorter than about a day no matter how fast the restore runs. Both terms are defined in business continuity standards such as ISO 22301.

Who should authorise DR invocation, and when?

Invocation should sit with a role that can accept the cost and the risk of failing over, typically an executive or the crisis lead, with a named deputy and a documented out-of-hours route. It is deliberately separated from the DR coordinator role in this chart: the coordinator assesses and recommends, management authorises, and the coordinator then runs the recovery. The trigger should be written in advance and be testable against facts available early in an outage, because the expensive mistake is not invoking too soon; it is spending three hours deciding whether this counts while the RTO clock runs.

What happens if the restored data fails its integrity check?

The chart loops rather than pushing on: 'Data integrity verified?' returns a failed branch to 'Restore from an earlier copy', which feeds back into the restore step so verification runs again. That loop is where the RPO gets tested in real conditions, because every step back to an older copy increases the data loss you will have to declare, and at some point the honest answer is that the target cannot be met and the business needs to be told what period it must reconstruct manually. It is worth agreeing in advance how many attempts you make, who is allowed to accept a recovery point worse than the RPO, and how the gap is recorded.

How often should the disaster recovery process be tested?

Often enough that the plan reflects the current estate, and with the results written down. Most organisations run a mix: frequent restore checks on backups, component or partial failover tests through the year, and a fuller exercise on a regular cycle, commonly annual, with regulated sectors and critical services expecting more. What matters more than the interval is what gets tested. A restore that is never actually mounted and read proves nothing, and an exercise that skips the validation and failback steps tends to hide the two problems that hurt most in a real event: application owners finding faults nobody anticipated, and no agreed route back to the primary site.

Use this template

More in IT process templates

More in Process map templates

Browse all IT process templates