incident-response.md
markdown
sha256:c2dbf04d56308f3bbf2d06e6d2eb022b8948b1e827195fe525a44e5e18d5f9c0
feat(auth): Phase B Connect cloud agent (RFC 8628) + Hermes…
Human
minor
⚠ breaking
12 days ago
title: Production incident response — SOP project: business-ops-template tags:
- playbook
- incident
- sre
- on-call date: 2026-04-07
Production incident response — SOP
Applies to: Customer-facing production services | Owner: Platform lead
Severity guide (use one)
- SEV1: Full outage or data loss risk; wake secondary on-call.
- SEV2: Major degradation; work business hours unless revenue-critical.
- SEV3: Minor issue with workaround; ticket and batch fix.
Immediate steps (first 15 minutes)
- Declare incident in the status tool; set severity; post customer-facing banner only if user-visible.
- Assign Incident Commander (IC) and scribe; IC drives, scribe timestamps actions.
- Capture symptoms: error rates, regions, last deploy, dependency status—link dashboards in the incident doc.
Stabilize before root cause
- Prefer rollback or feature flag off if change correlated; avoid speculative hotfixes during SEV1.
Communication
- Internal updates every 30 minutes for SEV1 until resolved.
- Customer comms go through support lead; no individual engineer posts externally.
After resolution & escalation
Post summary (window, cause, remediation, follow-ups in 48h); book a blameless retro within five business days with owned actions. If IC is absent, secondary runs IC; loop legal/comms on data exposure suspicion. Replace tool names and contacts for your org.
File History
1 commit
sha256:93bcf8f9bd56d8c5b9339f4ec73b9ebd66571398d56262d38eedc2cfa9db9882
fix(test): align Band B landing assertion with desktop MCP …
Human
12 days ago