Production Ops
production-ops owns:
How do we know production is healthy, and how do we recover when it is not?
Use when
Section titled “Use when”- a service is entering production
- runtime health signals change materially
- monitoring or alerting needs design
- incident response is active
- restore/recovery behavior needs proof
- capacity or operational toil is the problem
Don’t use when
Section titled “Don’t use when”Do not run an operations review for every code change.
runtime ↓health evidence ↓actionable event ↓safe response ↓recovery ↓recovery verificationTry it
Section titled “Try it”Use $production-ops to define health signals and recoveryfor the booking service after the new database dependency.