Skip to content
Talk to our solutions team

Run a disaster-recovery exercise

A rehearsed recovery: a simulated failure, recovered against the clock, with a per-phase record of how long each step took.

A completed exercise gives you a dated measurement of your recovery objectives and a list of procedural improvements to apply.

Record what the exercise covers before starting.

Exercise: DR-2026-09
Scope: loss of the primary data host for product <name>, staging environment
Out of scope: identity provider, network path, object storage
Success: service restored and verified within the stated RTO
Abort if: any production tenant is affected, or elapsed time reaches 2x RTO
Roles: who declares, who performs recovery, who observes and records timings

Define the abort condition explicitly. It keeps a rehearsal bounded and gives the observer clear authority to stop it.

Inform everyone who receives alerts for the target environment, including on-call. Alerts raised during the exercise are expected, and prior notice keeps them distinguishable from live incidents.

Inject the failure rather than describing it. Stopping the service is usually sufficient and is reversible.

Terminal window
kis flow -f dr-exercise.yaml -t simulate-loss -v target=staging-db-01

Start timing at the point of injection. Detection time is one of the figures being measured, so the clock starts before anyone is aware of the fault.

Follow the documented procedure as written.

PhaseFromToRecord
Detectfailure injectedfault observedWhether alerting fired or a person noticed
Decidefault observedrecovery declaredWho declared it and against what threshold
Provisiondeclaredcapacity availableWhich steps were manual
Restorecapacity availabledata loadedUse Restore from a backup
Verifydata loadedcorrectness confirmedData read back, not only a health probe
Cut overconfirmedserving trafficDNS, gateway routing, or standby promotion

Per-phase timings identify which step to improve, a single total figure confirms the outcome without indicating where the time was spent.

A readiness probe confirms the process started. Confirm the data separately.

Terminal window
kis flow -f dr-exercise.yaml -t verify-recovery -v target=staging

Check a value you can predict independently, such as a row count or a known record. This distinguishes a complete restore from one that produced a structurally valid but empty datastore.

Exercise DR-2026-09 15 September 2026
Scope: primary data host loss, staging
Detect 4m alert fired and was acknowledged
Decide 6m
Provision 18m
Restore 22m 40GB
Verify 15m
Cut over 12m
Total 1h17m against a 90m objective
Improvements
1 Pre-seed the image cache on replacement hosts to shorten provisioning.
2 Runbook credential reference updated to the current secret name.
3 Declaring role named explicitly in the runbook.
Next exercise: extend scope to include the gateway path.

Record the improvements alongside the timings. They are the output that changes the next result.

Action each item, then re-derive the objectives from the new measurement so the published figures match current behaviour.

ExerciseFrequency
Full exerciseTwice a year, and after any topology change
Restore-only rehearsalMonthly, against a scratch target
Confirmation runAfter remediation following a live incident

The monthly restore rehearsal is inexpensive, needs no announcement, and confirms that the backup chain is still producing usable archives.