Skip to content

Disaster recovery

Recovery objectives

ObjectiveTarget
RPO (data loss window)≤ 5 minutes for metadata (DynamoDB PITR), zero for document bodies (S3 versioning: every revision is a version)
RTO (time to restore)< 1 hour for full service (infra-as-code redeploy + data already durable)

What backs up what

DataMechanismRetention
Document metadata (DynamoDB)Point-in-time recovery, continuous35 days
Document bodies (S3)Versioning (every committed revision) + lifecycle → IA after 1 year, expire after 55 years
Static sites (S3)Source of truth is git + the pipeline; redeployable instantlyn/a
Infra definitionsGit (CDK) — every environment reproducible from a commitforever
Audit logsCloudTrail S3 (versioned, KMS) + CloudWatch Logs (1 year)1–5 years

Restore procedures

Restore a document body

Terminal window
# Latest committed revision is always the current object; older revisions:
aws s3api list-object-versions --bucket <bucket> --prefix private/<owner>/<id>/
aws s3api get-object --bucket <bucket> --key private/<owner>/<id>/<n>.frond \
--version-id <version> restored.frond

Restore metadata (DynamoDB PITR)

Console → DynamoDB → Backups → Restore to point-in-time (new table) → update the API’s DOCUMENTS_TABLE env var (CDK redeploy) or copy items back.

Full-environment rebuild

  1. git checkout <last-good-commit>
  2. Redeploy the environment (pipeline or cdk deploy -c env=<env> --all)
  3. Re-point DNS if the zone was lost (delegation records in the parent zone)
  4. Verify with the health endpoint (GET https://api.<domain>/healthz) and a smoke test (create/draw/save/load a canvas).

Region loss (extremely unlikely)

All state is in one region (us-east-1). A full regional outage is addressed by waiting for AWS recovery (RTO within AWS’s regional SLA) — acceptable for this risk tier; multi-region replication is a roadmap item if RTO targets change.

Testing

DR is exercised twice a year: one metadata PITR restore drill and one full-environment rebuild drill in dev.