test-engineering
Generated at build time from
skills/test-engineering/SKILL.md. Do not edit this page directly.
Testing exists to reduce uncertainty about meaningful behavior.
R → Graph → Proof<P, B, O>│ │ │ │ ││ │ │ │ └─ observable outcome / invariant (§6)│ │ │ └──── faithful test boundary (§3)│ │ └─────── property that must be proven (§2)│ ││ └─ nodes = setup, stimulus, system boundary, observations│└─ behavior / invariant / failure / regression risk§1 Gate testing itself must require design§2 P property, invariant, or regression to prove§3 B smallest faithful boundary where correctness lives§4 Setup deterministic state and controlled dependencies§5 Exercise stimulate the real behavior, including failures/races§6 O observe outcome, state, side effect, or invariant§7 Doubles isolate boundaries; never fake the semantics under proof§8 Regression reproduce the bug at the lowest useful faithful boundary§9 Suite reliability, speed, brittleness, and architecture§10 Handoff tests are executable proof; no testing ceremony artifactRead the behavior or risk. Decide exactly what must be proven. Find where that correctness actually lives. Use the smallest test that still exercises those semantics faithfully.
If a test does not materially increase confidence in a behavior, invariant, boundary, or failure mode, do not create it.
1. Applicability gate
Section titled “1. Applicability gate”test-engineering is a specialist capability, not a mandatory stage.
Do not invoke it when the verification is obvious:
simple deterministic business rule→ write a unit test directly
simple bug with an obvious reproduction→ write the regression test directly
existing local test pattern clearly fits→ follow it directly
mechanical/refactor-only change→ run targeted existing verificationInvoke test-engineering when testing itself has a design question, for example:
Which boundary should prove this behavior?
Unit, integration, contract, concurrency, or E2E?
How do we reproduce this race?
How do we prove transaction atomicity?
How do we verify retry / idempotency?
How do we test async or queue semantics?
How do we verify backward-compatible migration?
How do we inject a realistic dependency failure?
How do we prove backup/restore or recovery behavior?
Why is this suite flaky, slow, brittle, or expensive to maintain?Routing rule:
obvious verification→ implement directly
verification itself is complex or high-risk→ test-engineeringTesting is continuous feedback during implementation, not a phase after “dev complete.”
2. Define P: what exactly must be proven?
Section titled “2. Define P: what exactly must be proven?”Start from behavior and risk, not from files or functions.
Ask:
What can be wrong?What property would distinguish correct from incorrect?What observable evidence would prove it?Good proof targets include:
business rulecalculationparser/validator semanticsstate transitionpermission decisioninvariantpublic/API contracttransaction atomicityconcurrency outcomeidempotencyorderingretry semanticsmigration compatibilityfailure/recovery behaviorcritical user journeybug regressionExamples:
P1:two overlapping active bookings for the same room/timemust never both commit
P2:a retried payment command must not create a second charge
P3:records written by schema version N+1must still be readable by version N during rolloutDo not define the proof as:
"method X was called""mock Y received argument Z""line 42 executed""coverage increased"unless that call/line is itself part of a real external contract.
Test behavior, not implementation structure.
3. Choose B: the smallest faithful boundary
Section titled “3. Choose B: the smallest faithful boundary”Ask:
Where does correctness actually live?Then choose the smallest boundary that still contains those semantics.
Use when correctness is local and does not depend on real infrastructure semantics.
Strong candidates:
business rulescalculationsparsersvalidatorsstate machinespermission/policy decisionserror mappingretry decisionspure transformationsalgorithmic logicExample:
requested slot + existing slots ↓overlap policy ↓allowed / deniedA unit test is valuable because the rule itself is local.
Integration
Section titled “Integration”Use when correctness depends on a real boundary or infrastructure semantics:
SQL constrainttransactionlocking/isolationORM mappingfilesystem behaviorserializationrequest pipelinereal dependency protocoldatabase migrationDo not replace the semantics under proof with mocks.
If the risk is:
"only one concurrent booking may commit"the faithful boundary may require:
real concurrent execution+real transaction/storage boundaryA mocked transaction() call does not prove atomicity.
Contract
Section titled “Contract”Use when the primary risk is compatibility between independently evolving producer/consumer boundaries.
Examples:
API request/response shapeevent schemaconsumer/provider expectationbackward/forward compatibilityContract tests do not replace integration tests when runtime infrastructure semantics are the thing at risk.
Concurrency
Section titled “Concurrency”Use when the property only exists under overlapping execution.
Model the race explicitly:
start Astart Bsynchronize at critical pointrelease both ↓observe committed results ↓assert invariantAvoid “concurrency tests” that merely call the same function twice sequentially.
E2E / Critical User Journey
Section titled “E2E / Critical User Journey”Use when the property depends on the whole user-visible journey:
signup → verify → loginbooking → pay → confirmationedit → save → refresh → persisted stateKeep E2E count small and focused on critical journeys and failures that cannot be proven faithfully lower in the stack.
Specialized
Section titled “Specialized”Use specialized tests only when the risk demands them:
load / stressfault injectionproperty-basedfuzzcompatibilityrestore/recoverybrowser/network lifecycleDo not force them onto ordinary feature work.
Boundary rule
Section titled “Boundary rule”Choose the smallest testthat faithfully provesthe actual risk.“Smallest” without “faithful” creates false confidence.
“Faithful” without “smallest” creates slow, brittle suites.
4. Build deterministic setup
Section titled “4. Build deterministic setup”A test must control the preconditions needed for the proof.
Define:
initial statetest dataclock/timerandomnessidentityconfigurationdependency behaviorresource lifecycleconcurrency coordinationControl only what affects the property.
Prefer:
ephemeral dataisolated namespaces/databasesexplicit clocksseeded randomnessdeterministic synchronizationknown dependency fixturesAvoid:
shared mutable test statearbitrary sleepsdependency on test execution orderproduction data assumptionsuncontrolled wall clocknetwork timing as synchronizationIf the test cannot reproduce its own starting state, it is not reliable evidence.
5. Exercise the real behavior
Section titled “5. Exercise the real behavior”Stimulate the system through the boundary that owns the property.
Examples:
domain functionAPIdatabase transactionevent consumerworkerbrowserrestore procedureDo not reach into private helpers merely because they are easy to call.
For failure testing, inject the failure at the real dependency boundary:
payment provider timeoutdatabase serialization failurelost response after successful commitduplicate message deliveryworker restartpartial migration stateDo not simulate a failure in a way that bypasses the code path being validated.
For async systems, define completion from the behavior, not a sleep:
wait until observable state reaches Xoruntil bounded deadline expires6. Observe O: assert the property, not the choreography
Section titled “6. Observe O: assert the property, not the choreography”Assert on the smallest set of outputs that proves P.
Possible evidence:
return valuepublic responsepersisted stateevent emittedevent not duplicatedstate transitioninvariantexternal side effectdurable workvisible user outcomerecovered dataExample:
two concurrent booking attempts ↓assert:- exactly one succeeds- exactly one is rejected/retried as designed- one valid booking exists- overlap invariant still holdsAvoid assertions such as:
internal helper called onceexact private method orderingevery mock interactionincidental JSON field orderingexact UI text when wording is irrelevantA refactor that preserves behavior should not break the test.
7. Use doubles without invalidating the proof
Section titled “7. Use doubles without invalidating the proof”Rule:
Mock a boundary to isolate behavior.
Do not mock the thing whose semanticsyou are trying to prove.Good:
testing payment failure handling→ payment provider double can produce timeout/declineBad:
testing database transaction atomicity→ mock database transactionBad:
testing serialization compatibility→ bypass the real serializer with hand-built objectsChoose among:
fakestubmockreal dependencyephemeral dependencybased on the semantic fidelity required by P.
Prefer state/result assertions over interaction assertions.
Use mocks for interaction contracts only when the interaction itself is the behavior.
8. Regression rule
Section titled “8. Regression rule”For a real bug:
Bug ↓Can it be reproduced automaticallyat a useful faithful boundary? ↓yes→ create the smallest regression proof→ then fix / verify the fixThe test should fail for the original defect and pass for the corrected behavior.
Choose the boundary where the bug actually exists.
Example:
server commits completionbut response is lost ↓client refresh shows completed state ↓UI must not report false failureIf that behavior can be proven at service/integration level, do not require browser E2E.
If the defect only emerges from browser/network lifecycle, use E2E.
Do not write a regression test that merely repeats the implementation logic and therefore could share the same bug.
9. Special proof patterns
Section titled “9. Special proof patterns”Transactions and concurrency
Section titled “Transactions and concurrency”When correctness depends on isolation, locking, constraints, or atomicity:
real storage semantics+overlapping execution+final invariantDo not infer race safety from unit tests.
Test expected conflict/retry outcomes when the storage model can legitimately abort concurrent work.
Idempotency / retries
Section titled “Idempotency / retries”Model at least:
first deliveryduplicate deliveryretry after known failureretry after uncertain outcomeAssert on the final external effect:
one chargeone bookingone durable jobnot merely on handler return values.
Async / queue systems
Section titled “Async / queue systems”Consider only semantics that matter:
at-least-once deliveryduplicate messageorderingpoison messageretry/dead-letter behaviorcrash between effect and acknowledgementbacklog recoveryTest the property promised by the system; do not attempt to prove delivery guarantees the underlying broker does not provide.
Migrations / compatibility
Section titled “Migrations / compatibility”When old and new versions may coexist, test the actual compatibility window:
old schema/data ↓ migrate/expandold reader + new writer?new reader + old writer? ↓contract still valid ↓ contract/cleanupOnly test coexistence directions that the release strategy actually requires.
Use real migration scripts against representative schema/data where migration semantics are the risk.
Failure / recovery
Section titled “Failure / recovery”When the system claims it can recover:
establish valid state→ induce relevant failure→ execute recovery path→ assert state/invariants→ re-run critical behaviorA recovery command exiting 0 is not proof of recovery.
Security properties
Section titled “Security properties”security-review owns which attack/trust property matters.
test-engineering owns the executable proof when testing that property is non-trivial.
Example:
security-review:tenant A must never read tenant B data
test-engineering:design the faithful API/DB proof for that isolation propertyDo not independently invent a full security test matrix.
10. No coverage theater
Section titled “10. No coverage theater”Coverage is execution evidence, not correctness evidence.
Do not do:
coverage = 74%→ generate tests until 90%Use coverage only as a diagnostic clue:
important behavior appears unexercised ↓inspect whether a meaningful proof is missingDo not optimize line, branch, or mutation score without tying the added test back to behavior/risk.
A high-coverage suite can still fail to prove the important thing.
A low-coverage subsystem can still be adequately protected if most code is wiring and the meaningful behaviors are proven.
11. Treat flaky and brittle tests as defects
Section titled “11. Treat flaky and brittle tests as defects”A useful automated test should provide a trustworthy signal.
If the same code can randomly produce pass/fail, the test is not reliable evidence.
When diagnosing flakiness, inspect:
test setup/dataordering dependenceshared stateclock/timerandomnessasync synchronizationresource cleanuptest runner/frameworkSUT nondeterminismOS/network/external dependencyDo not normalize:
"rerun until green""known flaky""ignore intermittent failure"Short-term quarantine may protect the signal of the main suite, but it must not become permanent acceptance.
Brittleness is different from flakiness:
flaky→ same behavior, inconsistent result
brittle→ harmless implementation change breaks the testFor brittle tests, reduce knowledge of implementation details and assert on stable behavior.
12. Keep feedback proportional and fast
Section titled “12. Keep feedback proportional and fast”Fast feedback matters, but test fidelity comes first.
Structure the suite so the most relevant cheap proofs run earliest:
local targeted test ↓broader integration / contract ↓critical E2E / specialist suiteDo not run every expensive suite for every local edit when dependency analysis or CI stages can select them safely.
Do not split tests merely to hit a timing quota.
As a practical suite-health signal, automated developer/CI feedback should normally remain in minutes, not become a long stabilization phase. If the default suite becomes slow enough that developers avoid running it, its architecture needs work.
Separate expensive specialist suites when necessary:
loadlong-running compatibilitychaos/faultlarge E2Erecovery drillsbut ensure they run at the release/risk point where their evidence is needed.
13. Improve test architecture only when it pays
Section titled “13. Improve test architecture only when it pays”Use test-engineering for suite-level problems when they create real engineering cost:
flakinessslow feedbackduplicated setupbrittle selectors/contractsunrepresentative mockshard-to-create test dataunclear test ownershipexpensive environment boottests that cannot reproduce production bugsPrefer fixing the test boundary or system testability over adding more helpers around a bad design.
Examples:
hard-to-test time logic→ explicit clock boundary
massive mock graph→ component owns too many dependencies
every UI change breaks E2E→ assertions/selectors coupled to presentation details
DB bug only tested with mocks→ add focused real-DB integration boundaryDo not introduce a testing framework/abstraction unless it removes repeated complexity or enables a proof that was previously unreliable.
14. Relationship to other skills
Section titled “14. Relationship to other skills”design-thinking
Section titled “design-thinking”design-thinking owns obvious verification while implementing:
design→ implement→ obvious targeted proofHand off to test-engineering only when verification itself requires non-trivial boundary, environment, concurrency, failure, or suite design.
test-engineering may return a test design that design-thinking implements with the production change.
code-review
Section titled “code-review”code-review asks:
Does the implementation have meaningful risk?Is the existing verification faithful and sufficient?If the missing/incorrect verification is non-trivial:
code-review→ test-engineeringtest-engineering does not duplicate structural review.
engineering-design
Section titled “engineering-design”Engineering Design owns system guarantees and invariants.
Test Engineering turns those guarantees into executable proof when the test strategy is non-obvious.
If a guarantee cannot be tested without violating boundaries or massive mocking, that can reveal an engineering-design/testability problem.
security-review
Section titled “security-review”Security Review identifies security properties and attack paths.
Test Engineering designs reliable executable proof for those properties when needed.
production-ops
Section titled “production-ops”Production Ops owns what operational recovery/health must be proven.
Test Engineering can design complex recovery tests:
backup→ isolated restore→ integrity checks→ application smokeonly when that proof is non-trivial.
release-engineering
Section titled “release-engineering”Release Engineering decides which already-defined proofs become release gates and where they run.
Test Engineering designs the proof itself.
Do not put CI/CD orchestration in this skill.
15. Output and persistence
Section titled “15. Output and persistence”The primary output is executable tests and the minimum supporting fixtures/helpers needed to make those tests reliable.
Do not create:
.agents/tests.mdTEST_PLAN.mdQA_PLAN.mdcoverage-goal.mdby default.
The tests are the durable executable proof.
If a test strategy introduces a consequential architecture decision, such as a new system boundary solely to make critical behavior testable, feed that back to engineering-design rather than hiding the decision in test code.
For an analysis-only request, output:
Property→ risk to prove
Boundary→ why this is the smallest faithful test level
Setup→ required deterministic state/dependencies
Exercise→ action/failure/concurrency schedule
Assertions→ observable evidence
Gaps→ only material limitationsFor an implementation task:
design proof→ implement test→ run it→ confirm it can fail for the intended defect/property→ run relevant surrounding testsDo not report a test as proof merely because it passes once.
Proportionality
Section titled “Proportionality”Depth follows testing complexity and risk.
tiny deterministic logic→ direct unit/regression test; no specialist skill
normal feature→ targeted unit/integration tests using existing patterns
boundary/state/failure-sensitive feature→ test-engineering on affected proof graph
concurrency/migration/async/recovery/complex regression→ specialist test design + faithful environment
large or unhealthy suite→ test architecture analysis only where feedback/reliability cost justifies itDo not maximize test count.
Maximize:
confidence gained────────────────cost + brittleness + runtimesubject to the proof still being faithful.
The Pipeline
Section titled “The Pipeline”BEHAVIOR / RISK → "What exactly can be wrong?" → define P: property/invariant/regression
→ "Where does that correctness actually live?" → choose B: smallest faithful boundary
→ "What state and dependencies must be controlled?" → deterministic setup
→ "How do we exercise the real behavior?" → stimulate success/failure/race/compatibility path
→ "What observable outcome proves correctness?" → define O: result/state/invariant/side effect
→ "Am I mocking the semantics I am trying to prove?" → replace with real/ephemeral dependency when necessary
→ "Can this fail deterministically for the bug/property?" → validate test efficacy
→ "Does it remain useful after harmless refactors?" → remove implementation-detail coupling
→ EXECUTABLE PROOF → run at the cheapest lifecycle point where it preserves confidenceChoose the smallest test that faithfully proves the actual risk.