production-ops
Generated at build time from
skills/production-ops/SKILL.md. Do not edit this page directly.
Production operations is not “add monitoring.” It is the ability to know whether the workload is healthy, understand meaningful degradation, restore service or data when it is not, and prove recovery.
W → Graph → Ops<H, S, R>│ │ │ │ ││ │ │ │ └─ recovery paths that restore the workload (§6)│ │ │ └──── signals that detect and diagnose health (§3)│ │ └─────── measurable health of critical flows (§2)│ ││ └─ nodes = critical flows, dependencies, stateful resources│ edges = runtime dependencies and failure propagation│└─ production-sensitive workload or change§1 Context critical flows, runtime dependencies, state, risk§2 H healthy / degraded / unhealthy states§3 S detection signals first, diagnostic signals second§4 Alert page/notify only when a human action is required§5 Diagnose smallest evidence path from symptom to cause§6 R mitigation, recovery, and recovery verification§7 Data backup / restore / RPO / RTO only when state requires it§8 Readiness prove important operational paths before production§9 Incident assess → mitigate → recover → learn§10 Contract keep .agents/operations.md as a compact operational indexRead the workload or incident. Model the critical runtime flow. Decide what healthy means. Detect failure from symptoms users or critical flows experience. Add only the diagnosis and recovery machinery needed for the real risks.
If an operational control does not help detect, diagnose, mitigate, recover, or prove recovery from a material failure, remove it.
1. Establish the operational context
Section titled “1. Establish the operational context”Start from the running behavior and the failures that matter.
When present, read:
.agents/product.md.agents/engineering.mdrelevant ADRsruntime/deployment configurationexisting metrics/logs/tracesexisting alertsexisting backup/recovery configurationexisting runbooksrecent incident evidenceIdentify only what matters operationally:
- Critical flows — user/system outcomes whose failure makes the workload meaningfully degraded.
- Runtime dependencies — services, stores, queues, workers, external systems, network paths, or infrastructure required by those flows.
- Stateful resources — state whose loss, corruption, or unavailability changes recovery.
- Failure propagation — how one unhealthy node affects downstream outcomes.
- Recovery constraints — known acceptable downtime/data-loss limits, if they actually exist.
- Operational owner — who can respond, when that matters for alerting/escalation.
Do not invent scale, SLOs, RTO, RPO, pager rotations, retention periods, thresholds, or business impact.
For an existing workload change, model only the affected operational subgraph.
For a small personal workload, the graph may be:
request → application → databaseThat does not justify an enterprise observability stack.
2. Define H: health from critical outcomes
Section titled “2. Define H: health from critical outcomes”Health is not “CPU is below 80%.” Health is whether the critical workload behavior is still producing acceptable outcomes.
Model:
Critical flow ↓ depends onRuntime nodes ↓ produceObservable outcome ↓ classified ashealthy / degraded / unhealthyStart from the outside:
Can the user/system complete the critical operation?Is the result correct?Is it arriving within any known acceptable bound?Is accepted work durable when durability is required?Then derive the internal conditions that explain that outcome.
Example:
Checkout ├─ healthy → accepted orders complete correctly ├─ degraded → orders complete but latency/retry rate is abnormal └─ unhealthy → orders cannot be completed or correctness is at riskUse only health states the workload can actually distinguish and act on.
Do not create a formal SLO/error-budget program unless the workload, organization, or task needs one. If service objectives already exist, use them as health criteria rather than inventing parallel thresholds.
3. Derive S: detect symptoms before diagnosing causes
Section titled “3. Derive S: detect symptoms before diagnosing causes”Separate two questions:
What is broken?Why is it broken?Detection signals
Section titled “Detection signals”Prefer signals close to the critical outcome.
Typical candidates when relevant:
success / error ratelatencythroughput / trafficqueue age or backlogsaturationcorrectness invariant violationsfreshness / stalenessjob completiondependency availability visible at the workload boundaryFor user-facing request/response systems, latency, traffic, errors, and saturation are useful defaults to consider, not mandatory metrics.
For workers or pipelines, better primary signals may be:
accepted workcompleted workfailure rateoldest work agebacklogprocessing latencyFor data systems:
availabilitywrite/read correctnessreplication or freshness lagcapacityrecovery statusDiagnostic signals
Section titled “Diagnostic signals”Add only the evidence needed to explain a detected symptom:
structured logstracesdependency metricsresource metricsstate transition recordsdeployment/version metadatacorrelation/request IDsDo not make verbose logs the primary health detector when a direct outcome signal exists.
Do not instrument every function. Instrument boundaries, critical transitions, and failure points that materially shorten diagnosis.
Telemetry exists to answer an operational question. If a signal has no question, owner, or use, do not add it.
4. Alert only when action is required
Section titled “4. Alert only when action is required”An alert is an interrupt. It needs a reason.
Create an alert only when:
a material health condition exists+someone can take a meaningful action+waiting for passive inspection would materially worsen the outcomeAlert primarily on symptoms or direct health-state changes.
Use cause-level alerts only when the cause itself requires independent immediate action.
Every actionable alert should make these clear:
What is affected?How bad is it?Where should the responder look first?What safe first action or runbook is available?Avoid:
one alert per metricalerts with no responderalerts that have no actionduplicate alerts for the same user-visible symptomstatic thresholds copied from generic guidanceNot every abnormal signal needs an alert.
Depending on risk, it may be enough to:
record itshow it on demandsurface it during diagnosisChoose paging, asynchronous notification, dashboard visibility, or logs according to urgency and required human action.
5. Build the shortest diagnosis path
Section titled “5. Build the shortest diagnosis path”Once a symptom is detected, a responder should be able to traverse:
health symptom ↓affected critical flow ↓recent change / dependency / state ↓evidence ↓likely failure domainPrefer a small number of views organized around critical flows over dashboards organized around every available metric.
A useful diagnostic view answers questions such as:
What changed?Which flow is unhealthy?When did it start?Which dependency or state transition correlates?Is the failure global or scoped?Is the workload failing, slow, saturated, stale, or incorrect?Use logs, traces, dashboards, queries, or runtime inspection only where each gives evidence the others cannot provide cheaply.
Do not build a dashboard as proof of observability.
The proof is that a realistic failure can be detected and narrowed to an actionable failure domain.
6. Define R: mitigation, recovery, and proof
Section titled “6. Define R: mitigation, recovery, and proof”For every material failure class, ask:
How do we stop user impact?How do we restore a valid workload state?How do we know recovery actually worked?Separate:
Mitigation
Section titled “Mitigation”Reduce impact before full root cause is known.
Examples when applicable:
disable a broken feature pathshed or reroute trafficpause a workerisolate a bad dependencyfail overrestart a stuck componentstop a destructive jobrequest rollback of a bad releaseRecovery
Section titled “Recovery”Return the workload to a valid operating state.
Examples:
restart/recreate state safelyreprocess durable workrestore datarepair corrupted stateroll forward with a fixuse release-engineering rollbackfail back after dependency recoveryVerification
Section titled “Verification”Never infer recovery from “the command succeeded.”
Re-check the original health outcome:
mitigation/recovery action ↓critical flow recovers? ↓state/invariant valid? ↓backlog/drain/freshness normal? ↓no continuing user impact?Release-engineering boundary
Section titled “Release-engineering boundary”production-ops owns:
detect degradationassess impactdetermine that mitigation/recovery is neededselect the operational objectiveverify health after recoveryrelease-engineering owns:
deploymentrollout mechanicsrollback / roll-forward mechanismartifact/environment promotionExample:
new release causes elevated failures ↓production-ops:"current production health is unacceptable" ↓release-engineering:execute the defined rollback mechanism ↓production-ops:verify the critical flow is healthy againDo not redesign deployment machinery inside Production Ops.
7. Treat data recovery as a proven capability
Section titled “7. Treat data recovery as a proven capability”A backup file is not recovery.
For state whose loss matters:
state ↓acceptable data loss? → RPO, only if required/knownacceptable downtime? → RTO, only if required/known ↓backup / replication / rebuild strategy ↓restore procedure ↓integrity / usability verificationDo not invent RTO or RPO.
If the workload can reconstruct data from another authoritative source, document and test that reconstruction instead of adding redundant backup machinery.
For backups that matter, prove:
backup existsbackup is accessible to the recovery processrestore succeedsrestored data is usable and internally validrecovery procedure is executable under realistic conditionsA backup strategy is incomplete until restore is tested.
Prefer recovery tests that use the same procedure expected during a real incident.
Do not require multi-region DR, replicas, point-in-time recovery, or automated failover unless the recovery requirements justify their cost and complexity.
8. Prove operational readiness before production when risk justifies it
Section titled “8. Prove operational readiness before production when risk justifies it”Production Ops applies before first launch and before changes that materially alter the operational envelope.
For the affected subgraph, ask:
Can we tell whether the critical flow is healthy?Can we detect its important failures?Can we diagnose the likely failure domain?Can we perform the required mitigation/recovery?Can we prove recovery?Can we recover important state if it is lost?Only add a runbook when a non-obvious human procedure is required.
Only add an alert when a human must be interrupted.
Only test recovery paths that matter to the workload’s actual failure risk.
Examples:
New stateless HTTP service
Section titled “New stateless HTTP service”Likely enough:
health / request success signalerror + latency visibilitydependency failure visibilityrestart/redeploy recoverysmoke/recovery verificationNew durable background worker
Section titled “New durable background worker”Likely needs:
accepted/completed work signalfailure ratebacklog / oldest job ageretry/idempotency behaviorstuck-worker diagnosisrestart/replay recoveryproof that durable work is not lostStateful production database change
Section titled “Stateful production database change”May require:
data-health signalcapacity/failure visibilitybackup/restore proofmigration/recovery compatibilityexplicit recovery constraintsDo not expand one affected subgraph into a company-wide reliability program.
9. Handle incidents mitigation-first
Section titled “9. Handle incidents mitigation-first”During an active production incident:
detect ↓assess impact ↓stabilize / mitigate ↓recover service or state ↓verify health ↓diagnose root cause deeply ↓prevent recurrenceDo not delay a safe mitigation while searching for perfect root-cause certainty.
Preserve useful evidence when mitigation could destroy it, but user/service impact takes priority.
For small incidents, one responder may perform all roles. Do not impose incident-command bureaucracy where it adds no value.
For large incidents, separate coordination, technical operations, and communication when doing so reduces confusion.
During response, keep a concise timeline of:
symptomimpacthypothesesactionsobserved resultscurrent healthAfter meaningful incidents, capture only learning that changes the system:
detection gapdiagnosis gapmitigation gaprecovery gaparchitecture/control gapCreate follow-up work with a verifiable end state.
Do not turn every small production bug into a postmortem ceremony.
10. Persist only durable operational context
Section titled “10. Persist only durable operational context”When working in a project, maintain a compact:
.agents/operations.mdThis file is an index of durable operational truth, not the executable monitoring or deployment system.
Use the smallest useful structure:
# Operations
## Critical Health- <flow>: <healthy/degraded/unhealthy condition>
## Detection- <signal / alert>: <what it proves and when action is required>
## Diagnosis- <entry point to the evidence path>
## Recovery- <failure>: <mitigation/recovery path> → <recovery proof>
## Data Recovery- <state>: <backup/rebuild + restore proof>
## Runbooks- <only non-obvious human procedures>
## Risks / Assumptions- <only unresolved items that materially affect operations>Omit sections that do not apply.
Do not create separate observability, monitoring, alerting, reliability, backup, DR, or incident-management documents by default.
Executable configuration remains the source of truth for mechanics:
monitoring / alert rulesdashboards as code when usedapplication instrumentationbackup configurationruntime configurationautomationrelease configuration.agents/operations.md should point to or summarize those mechanics only when future operators/agents need the context.
If .agents/operations.md already exists:
- read it first;
- preserve still-valid operational decisions;
- update only the affected operational subgraph;
- remove stale references and contradictions;
- do not rewrite unrelated operations context.
Relationship to other skills
Section titled “Relationship to other skills”engineering-design
Section titled “engineering-design”engineering-design→ defines production-relevant system guarantees, state ownership, boundaries, and failure semantics
production-ops→ turns those guarantees into runtime health, detection, diagnosis, and recovery capabilityIf Production Ops discovers that recovery or health is impossible because of an architectural limitation, feed that requirement back to engineering-design.
release-engineering
Section titled “release-engineering”release-engineering→ how a change reaches or leaves production safely
production-ops→ whether production is healthy, whether recovery is required, and whether recovery succeededProduction Ops consumes release/version metadata for diagnosis but does not own the deployment pipeline.
design-thinking
Section titled “design-thinking”Use design-thinking when production-ops requirements require application changes such as:
instrument a boundaryadd a health endpointmake retry idempotentexpose a diagnostic stateimplement graceful recoveryProduction Ops defines the operational property; Design Thinking implements the software graph.
code-review
Section titled “code-review”Code Review checks whether operationally relevant changes actually preserve the intended health, recovery, and failure semantics.
Do not use Production Ops as a substitute for reviewing application correctness.
Proportionality
Section titled “Proportionality”Depth follows operational risk.
local/personal app→ basic failure visibility + local recovery + data restore if needed
small production service→ critical-flow health + actionable failures + recovery proof
stateful / async / externally dependent workload→ affected dependency/state subgraph + recovery readiness
high-impact distributed workload→ explicit health model + actionable alerting + tested recovery + incident/runbook support where justifiedDo not maximize monitoring coverage.
Maximize confidence that important production failures are:
detectable→ diagnosable enough→ recoverable→ verifiably recoveredThe Pipeline
Section titled “The Pipeline”WORKLOAD / CHANGE / INCIDENT → "Which critical runtime outcome matters?" → identify the affected operational subgraph
→ "What does healthy, degraded, and unhealthy mean?" → define H from observable outcomes
→ "What tells us WHAT is broken?" → derive detection signals
→ "What evidence tells us WHY?" → add only required diagnostic signals
→ "Does this condition require immediate human action?" → alert only when actionable
→ "How do we stop impact and restore valid state?" → define mitigation and recovery
→ "How do we prove recovery?" → re-check the original health outcome and invariants
→ "Can important state be recovered?" → test restore/rebuild when data risk requires it
→ before production-sensitive launch/change → prove only the affected readiness paths
→ during incident → assess → mitigate → recover → verify → learn
→ .agents/operations.md → compact operational context; executable config remains truthIf production can fail in a material way and you cannot detect it or prove recovery from it, the operational design is incomplete.