Run a disaster-recovery exercise
A rehearsed recovery: a simulated failure, recovered against the clock, with a per-phase record of how long each step took.
A completed exercise gives you a dated measurement of your recovery objectives and a list of procedural improvements to apply.
Step 1: define the scope
Section titled “Step 1: define the scope”Record what the exercise covers before starting.
Exercise: DR-2026-09Scope: loss of the primary data host for product <name>, staging environmentOut of scope: identity provider, network path, object storageSuccess: service restored and verified within the stated RTOAbort if: any production tenant is affected, or elapsed time reaches 2x RTORoles: who declares, who performs recovery, who observes and records timingsDefine the abort condition explicitly. It keeps a rehearsal bounded and gives the observer clear authority to stop it.
Step 2: notify
Section titled “Step 2: notify”Inform everyone who receives alerts for the target environment, including on-call. Alerts raised during the exercise are expected, and prior notice keeps them distinguishable from live incidents.
Step 3: simulate the failure
Section titled “Step 3: simulate the failure”Inject the failure rather than describing it. Stopping the service is usually sufficient and is reversible.
kis flow -f dr-exercise.yaml -t simulate-loss -v target=staging-db-01Start timing at the point of injection. Detection time is one of the figures being measured, so the clock starts before anyone is aware of the fault.
Step 4: recover, timing each phase
Section titled “Step 4: recover, timing each phase”Follow the documented procedure as written.
| Phase | From | To | Record |
|---|---|---|---|
| Detect | failure injected | fault observed | Whether alerting fired or a person noticed |
| Decide | fault observed | recovery declared | Who declared it and against what threshold |
| Provision | declared | capacity available | Which steps were manual |
| Restore | capacity available | data loaded | Use Restore from a backup |
| Verify | data loaded | correctness confirmed | Data read back, not only a health probe |
| Cut over | confirmed | serving traffic | DNS, gateway routing, or standby promotion |
Per-phase timings identify which step to improve, a single total figure confirms the outcome without indicating where the time was spent.
Step 5: verify with data
Section titled “Step 5: verify with data”A readiness probe confirms the process started. Confirm the data separately.
kis flow -f dr-exercise.yaml -t verify-recovery -v target=stagingCheck a value you can predict independently, such as a row count or a known record. This distinguishes a complete restore from one that produced a structurally valid but empty datastore.
Step 6: record the result
Section titled “Step 6: record the result”Exercise DR-2026-09 15 September 2026Scope: primary data host loss, staging
Detect 4m alert fired and was acknowledgedDecide 6mProvision 18mRestore 22m 40GBVerify 15mCut over 12mTotal 1h17m against a 90m objective
Improvements 1 Pre-seed the image cache on replacement hosts to shorten provisioning. 2 Runbook credential reference updated to the current secret name. 3 Declaring role named explicitly in the runbook.
Next exercise: extend scope to include the gateway path.Record the improvements alongside the timings. They are the output that changes the next result.
Step 7: apply the improvements
Section titled “Step 7: apply the improvements”Action each item, then re-derive the objectives from the new measurement so the published figures match current behaviour.
Cadence
Section titled “Cadence”| Exercise | Frequency |
|---|---|
| Full exercise | Twice a year, and after any topology change |
| Restore-only rehearsal | Monthly, against a scratch target |
| Confirmation run | After remediation following a live incident |
The monthly restore rehearsal is inexpensive, needs no announcement, and confirms that the backup chain is still producing usable archives.