Production incident
For an active incident:
detect → assess impact → stabilize → mitigate → verify recoveryUse production-ops for that path.
After stabilization:
implementation defect?→ design-thinking / code-review
trust or authorization issue?→ security-review
race / transaction / recovery proof is hard?→ test-engineering
system ownership or guarantee is wrong?→ engineering-designDo not delay mitigation while users are still affected.