Skip to content

Production Ops

production-ops owns:

How do we know production is healthy, and how do we recover when it is not?

  • a service is entering production
  • runtime health signals change materially
  • monitoring or alerting needs design
  • incident response is active
  • restore/recovery behavior needs proof
  • capacity or operational toil is the problem

Do not run an operations review for every code change.

runtime
↓
health evidence
↓
actionable event
↓
safe response
↓
recovery
↓
recovery verification
Use $production-ops to define health signals and recovery
for the booking service after the new database dependency.

Read the complete production-ops instructions.