Business Continuity & Disaster Recovery
How an organisation keeps working through a disruption and gets its systems back afterwards. Business continuity (BC) keeps critical business functions running during the event, perhaps on paper or at another site. Disaster recovery (DR) restores the IT systems and data those functions depend on. DR is a subset of BC.
Business Impact Analysis and Recovery Objectives
What is this?
A business impact analysis (BIA) asks each part of the business: if this process stopped, what would it cost per hour or per day, and how long could we survive without it? The answers become the recovery targets that every backup and DR design must meet.
Why it matters
Without a BIA, DR spending is guesswork: either gold-plating systems nobody needs fast, or discovering mid-ransomware that the payment system had a one-week restore time. Interviewers use RTO and RPO to test whether you can design for a requirement rather than for a technology.
How it works
Time ───────────────────────────────────────────────────────────────────────►
last good backup disaster service restored
│ │ │
│◄──────── RPO ────────►│◄──────────── RTO ─────────►│
│ data you can afford │ downtime you can afford │
│ to lose │ │
│◄────────────── MTD ─────────────────►│
beyond this, the business may not survivethe maximum acceptable data loss, measured in time. An RPO of 15 minutes means backups or replication at least every 15 minutes.
the maximum acceptable time to restore the service.
(also MAD, maximum allowable downtime): the point beyond which the damage to the business is unacceptable. RTO must be shorter than MTD, with margin for verification (sometimes called WRT, work recovery time: RTO + WRT ≤ MTD).
processes are ranked so the most critical recover first. Dependencies matter: the payment service is useless if identity and DNS are down.
BIA steps:
- Identify business processes and their owners.
- Map each process to the systems, people, suppliers and data it needs.
- Estimate the impact of an outage over time: financial, regulatory, reputational, safety.
- Set MTD, RTO and RPO for each process.
- Feed the results into the continuity strategy and the budget.
In practiceA BIA ends up as a table that engineering can design against:
| Process | Owner | Cost of outage | MTD | RTO | RPO | Depends on |
|---|---|---|---|---|---|---|
| Card payments | Head of Payments | £40k/hour + regulatory | 8 h | 2 h | 5 min | Identity, payments DB, card processor, DNS |
| Payroll | Finance Director | Low until payday, then severe | 3 days | 24 h | 24 h | HR system, bank file transfer |
| Marketing site | CMO | Reputational | 2 days | 12 h | 24 h | CDN, CMS |
The "depends on" column is where plans usually fail: payments can't recover in 2 hours if the identity provider it needs has a 24-hour RTO.
📰Real incident — Maersk and NotPetya (2017). Destructive malware wiped tens of thousands of Maersk machines in hours, including every online domain controller. The company recovered its Active Directory from a single domain controller in Ghana that happened to be offline during a power cut. Global shipping operations ran on manual workarounds for about ten days; the cost was estimated at around $300 million. Identity was the hidden dependency of everything.
🎯On the job. Ask every critical service owner two questions: "what's your real RTO, measured in the last test?" and "what do you depend on to come back?" The answers rarely match the documented plan.
Security angle
Security incidents are now the most likely "disaster." Ransomware recovery is a DR scenario, but with a twist: you must be sure backups are clean and the attacker is gone before restoring, or you restore the attacker too.
Backup Strategies
What is this?
The ways data is copied so it can be restored.
Why it matters
The backup type decides how long backups take, how much storage they need and, crucially, how long and complex a restore is.
How it works
| Type | What it copies | Backup time | Restore needs |
|---|---|---|---|
| Full | Everything | Longest | Just the last full |
| Incremental | Changes since the last backup of any kind; clears the archive bit | Shortest | Last full + every incremental since |
| Differential | Changes since the last full; does not clear the archive bit | Grows each day | Last full + latest differential only |
three copies of data, on two different media, one off-site. Modern versions add "1 immutable or offline copy and 0 errors in restore tests" (3-2-1-1-0).
near-instant copies and continuous replication give low RPOs, but replicate corruption and deletion too. They complement backups; they don't replace them.
write-once storage (object lock, vault lock) so an attacker with admin rights cannot delete backups. See AWS backups and ransomware resilience.
sends bulk backups off-site; remote journaling sends transaction logs continuously; database shadowing keeps a live remote copy.
In practice"Is this backup actually restorable and immutable?" is answerable from the command line:
$ restic -r s3:s3.amazonaws.com/acme-backups snapshots --latest 3
ID Time Host Paths
4f1c2a9e 2026-10-11 02:00:04 db01 /var/lib/postgresql ← last night's snapshot exists
$ restic -r s3:… check --read-data-subset=5% ← actually read data back, not just list it
no errors were found
$ aws s3api get-object-lock-configuration --bucket acme-backups
{ "ObjectLockConfiguration": { "ObjectLockEnabled": "Enabled",
"Rule": { "DefaultRetention": { "Mode": "COMPLIANCE", "Days": 35 } } } } ← nobody can delete for 35 days📰Real incident — GitLab database deletion (2017). An engineer accidentally deleted the production database directory while fixing replication. GitLab then discovered that none of its five backup and replication methods worked as expected: scheduled dumps had been silently failing. Recovery came from a staging snapshot taken six hours earlier, and about six hours of data was lost. They livestreamed the recovery and published a famously honest postmortem.
🎯On the job. Alert on backup failures and on backups that are suspiciously small, and run restore tests on a schedule. The only proof a backup works is a restore.
Security angle
Backups are a high-value target: they contain everything. Encrypt them, store them in a separate security domain with separate credentials, and alert on deletion attempts.
Recovery Sites and High Availability
What is this?
Where systems run when the primary site is unavailable, and how designs avoid single points of failure.
Why it matters
The recovery site choice is a cost-versus-RTO trade-off, and it is a classic exam question.
How it works
| Site type | What's there | Typical RTO | Cost |
|---|---|---|---|
| Hot site | Fully equipped, data near-current, ready to run | Minutes to hours | Highest |
| Warm site | Hardware and connectivity, data needs restoring | Hours to days | Medium |
| Cold site | Space, power, cooling; no equipment | Weeks | Lowest |
| Mobile site | Equipment in a trailer or container | Days | Varies |
| Reciprocal agreement | Another organisation shares its facility | Unreliable | Low |
| Redundant (mirrored) site | Second live site, active–active | Near zero | Very high |
Cloud equivalents (fastest to slowest recovery): multi-site active–active, warm standby (a scaled-down live copy), pilot light (core data replicated, servers off until needed), and backup and restore.
High availability building blocks: redundant components and power, clustering and load balancing, RAID (redundant array of independent disks — RAID 1 mirroring, RAID 5 striping with parity tolerates one disk failure, RAID 6 two, RAID 10 mirrored stripes), multiple availability zones or regions, and avoiding shared dependencies (one DNS provider, one identity provider).
📰Real incident — OVHcloud data-centre fire (2021). A fire destroyed one data centre and damaged another at OVHcloud's Strasbourg site. Customers who had kept their backups in the same site — sometimes the same building — lost data permanently. "Off-site" has to mean a different failure domain, not just a different server.
📰Real incident — CrowdStrike outage recovery (2024). When a faulty agent update crashed millions of Windows machines, many organisations had to fix each machine by hand in safe mode. Those using BitLocker disk encryption first needed each machine's recovery key — and some had stored those keys on systems that were themselves down. Recovery dependencies include the secrets you need to recover.
🎯On the job. Keep break-glass essentials — recovery keys, admin credentials, network diagrams, the DR runbook, contact lists — available offline or in a separate provider, because the incident may take down the place you normally keep them.
Security angle
A DR region must have the same security controls as production: logging, guardrails, secrets and detection. Attackers like DR environments because they are rarely watched.
Testing the Plan
What is this?
The different ways to prove a continuity or DR plan works, in order of increasing realism and risk.
Why it matters
An untested plan fails on first use: contact lists are stale, the runbook references a decommissioned server, the restore takes four times the RTO.
How it works
owners review the plan for accuracy.
the team talks through a scenario step by step around a table.
a realistic scenario is acted out, sometimes with partial technical steps, without failing over production.
the recovery site is brought up and processes real data alongside the primary, which keeps running.
the primary is actually shut down and operations move to the recovery site. Most realistic, highest risk; needs senior approval.
After every test or real event, run a lessons-learned review and update the plan. See postmortems.
In practiceA useful tabletop is a timed scenario with injects, not a slideshow:
09:00 Inject 1: Helpdesk reports staff can't open files; ransom note on file server.
09:15 Inject 2: Backup admin finds backup server console shows "repository deleted".
09:40 Inject 3: A journalist emails asking about "the ransomware attack on your company".
10:00 Inject 4: Cyber insurer asks whether you have engaged their panel responder yet.
Questions at each inject: Who decides? What do we isolate? Who do we tell, by when? What do we restore first?🎯On the job. Measure each exercise: time to declare an incident, time to first containment action, time to a working restore of one critical system. Track them across exercises like any other metric.
Security angle
Include a ransomware scenario: restore from immutable backups into an isolated environment, verify the data is clean, and time it against the RTO. That is the test most organisations have never run.
Interview Questions
RPO is how much data you can afford to lose, measured as time since the last good copy; RTO is how long you can afford to be down; and maximum tolerable downtime is the point after which the business is seriously harmed. RTO has to fit inside MTD with room to verify the restore. They come out of the business impact analysis, and they're what decide whether you need replication, a warm standby or plain backups.
Differential restores faster because you need only the last full backup plus the most recent differential, whereas incremental needs the full plus every incremental since, in order. Incrementals are quicker and smaller to take each night, so it's a trade between backup window and restore time. For ransomware resilience, what matters more is that at least one copy is immutable and restores are actually tested.
By the RTO the business impact analysis sets and what the business will pay. A hot site recovers in minutes to hours but costs nearly as much as production; a warm site has hardware but needs data restored, so hours to days; a cold site is just space and power, so weeks. In the cloud the same spectrum is active-active, warm standby, pilot light, and backup and restore.
Traditional DR assumes the disaster is done when you start recovering; ransomware assumes an attacker who may still be inside and who deliberately targets backups. So backups must be immutable and held in a separate security domain, restores happen into a clean, isolated environment, and you confirm the attacker's access and persistence are removed first. Otherwise you restore straight back into a compromised environment.
In order of realism: read-through, tabletop walkthrough, simulation, parallel test and full interruption. I'd start with a read-through and tabletop because they're cheap and catch stale contacts and missing steps, then move to a parallel test that proves the recovery site can process real data without touching production. Full interruption is the only real proof, but it needs senior sign-off because a failed test is a real outage.