Derive your recovery objectives
Two measured figures for your deployment: the maximum data loss a failure can cause (RPO) and the time to restore service (RTO).
Why they are deployment-specific
Section titled “Why they are deployment-specific”kis.ai runs on your infrastructure. Backup schedule, storage location, network and topology are yours to choose, and each one moves the result, so the objectives are derived from a specific deployment rather than fixed by the platform.
| Input | Affects |
|---|---|
| Backup frequency | RPO |
| Whether archives are stored off-host | RPO |
| Whether the datastore replicates | RPO |
| Restore duration at your data volume | RTO |
| Time to provision replacement capacity | RTO |
| Whether a standby is already running | RTO |
Record the derivation alongside the figures. A stated method lets the numbers be reproduced and re-checked as the deployment changes.
Step 1: calculate RPO
Section titled “Step 1: calculate RPO”Worst-case data loss is everything written since the last verified archive.
RPO = backup interval + detection timeDetection time belongs in the calculation. A nightly backup gives a 24 hour RPO only when a failure is noticed within the same cycle. If a fault goes unnoticed for two further days, the archives taken in between captured it, and the recoverable point is three days back.
Two levers shorten it:
| Lever | Effect |
|---|---|
| A shorter backup interval | Reduces the first term directly |
| Verification on every backup run | Reduces the second by surfacing a bad archive at the time it is written |
Back up a service covers a flow that verifies each archive before reporting success.
Step 2: measure RTO
Section titled “Step 2: measure RTO”Restore into a scratch target and time it.
time kis flow -f restore-service.yaml \ -v archive=backup-20260915T020000Z.tar.gz \ -v target=scratchRestore from a backup defaults to a scratch target and keeps the production path behind an explicit flag, so it is safe to run for measurement.
The restore is one phase of several:
RTO = detect + decide + provision + restore + verify + cut overMeasure each phase separately. Provisioning replacement capacity and the decision to declare a failure are frequently larger than the restore itself, and a single end-to-end figure hides which phase to work on.
Step 3: record the figures with their inputs
Section titled “Step 3: record the figures with their inputs”RPO: 6h backups every 4h, verified each run, detection within 2h via readiness alertingRTO: 90m restore of 40GB measured at 22m; provision 30m; verify 15m; cut over 20mMeasured: 2026-09-15, against the production data volumeRe-derive whenever an input changes. Data volume growth moves RTO without any configuration change, so a periodic re-measurement keeps the figure current.
Step 4: improve them
Section titled “Step 4: improve them”Recovery objectives follow from the topology. To change the figures, change the inputs:
| Target | Approach |
|---|---|
| Lower RPO | Increase backup frequency; replicate the datastore in addition to archiving |
| Lower RTO | Run a standby so provisioning is complete before a failure |
| Both | Rehearse the procedure, which shortens the detect and decide phases |
See HA deployment for standby topology and Run a disaster-recovery exercise for rehearsal.
Verify
Section titled “Verify”The objectives are reproducible when a second engineer can derive the same figures from the recorded inputs without consulting whoever measured them.