Skip to content

Production incident

For an active incident:

detect
→ assess impact
→ stabilize
→ mitigate
→ verify recovery

Use production-ops for that path.

After stabilization:

implementation defect?
→ design-thinking / code-review
trust or authorization issue?
→ security-review
race / transaction / recovery proof is hard?
→ test-engineering
system ownership or guarantee is wrong?
→ engineering-design

Do not delay mitigation while users are still affected.