Incident response
Severity levels
| Level | Definition | Response target |
|---|---|---|
| SEV1 | Data breach, data loss, service-wide outage | Page on-call; contain within 1h |
| SEV2 | Partial outage, security finding with limited impact | Within 4 business hours |
| SEV3 | Minor bug, single-user issue | Next business day |
1. Detect
Sources: CloudWatch alarms (SNS), GuardDuty findings, WAF anomalies, budget
alarms, customer reports via security@…, private vulnerability reports
(SECURITY.md).
2. Triage & declare
- Appoint an incident commander; create a channel/ticket with timestamps.
- Classify severity; log a timeline.
3. Contain
- Data exposure: disable publishing for affected documents (API toggle), rotate the KMS key material if warranted, revoke affected Cognito tokens (admin revoke), block offending IPs in WAF.
- Service outage: roll back the last deploy (pipeline history); frontends are immutable assets so rollback is instant.
- Compromised credentials: deactivate IAM keys/Cognito user, rotate.
4. Investigate
- CloudTrail (org trail) for the blast radius; S3 access logs; Lambda logs (never contain PHI); WAF logs.
- Determine: what was accessed, by whom, when, and whether PHI was involved (this decides whether the HIPAA 60-day notification clock applies).
5. Recover & notify
- Restore from S3 versions / DynamoDB PITR if needed (DR guide).
- Notify: internal stakeholders; affected users; regulators within legal deadlines (60 days for HIPAA breaches of unsecured PHI).
6. Post-incident review
- Timeline write-up, root cause, control gaps → tickets with owners.
- Update this runbook and the threat model.
Contacts
(Populated at launch with real names; roles: incident commander, security owner, comms owner, legal.)