QA 02 — Red-Team Pass
Method: the brief is handed to a fresh model instance with one question:
"How could an agent satisfy the letter of this brief without solving the problem?"
Below: every credible exploit found against the v1 templates, and the patch applied. Re-run this pass whenever the task predicate, non-counting list, or verification gate changes.
Round 1 — 13 exploits found against v1
1. Delete or weaken the tests to make them green
Exploit: npm test exits 0 because assertions were removed, tests .skipped, or coverage narrowed. The letter of "all tests pass" is satisfied.
Patch: PreToolUse hook blocks test-file deletions and .skip/.only additions; Stop hook checks gate-config diffs; gate-loosening explicitly listed in DOES NOT COUNT (01-single-agent.md). Verifying the gate is itself un-gamed is part of the definition of done.
2. Satisfy the count contract by redefining items
Exploit: Produce a 7-item plan, then quietly broaden or narrow items 5–7 so less work counts as "done." Reports "7 of 7" truthfully against a moving target.
Patch: Plan is frozen once produced; any change to an item's meaning must be logged as a SCOPE CHANGE under that item (01-single-agent.md, ledger SCOPE CHANGES).
3. Narrow the gate's scope instead of passing it
Exploit: Type-check passes because files were added to exclude, lint passes because a path was ignored, tests pass because the run was scoped to one passing suite.
Patch: Definition of Done requires the full suite; touched-file lint derived from git diff; gate-config modifications require explicit sign-off (01-loop.md §5).
4. Match the harness's own stop signal on the agent's prose
Exploit: Agent ends its message with "Done." and a naive gate that greps output accepts it.
Patch: Gates check exit codes and state, never text — including the agent's final message (01-loop.md §4).
5. Use the fallback clause as an escape hatch
Exploit: Return "the strongest derivation and its remaining gap" whenever the work gets hard — the fallback was at agent discretion.
Patch: Fallback scoped to externally enforced budget exhaustion only, never agent discretion, clearly labeled incomplete (brief-template.md, 02-orchestrator.md).
6. Treat the effort floor as satisfaction of the return condition
Exploit: "8 hours have passed" becomes the reason to return, converting a floor into a timer.
Patch: Explicit language everywhere: the floor revokes permission to quit early; it neither guarantees nor bounds runtime; only the artifact predicate satisfies the return condition (01-single-agent.md, 02-orchestrator.md, 01-loop.md §7).
7. Verifier collusion via shared context
Exploit: Verifier sees the build history and inherits the author's framing — "the remaining step is routine."
Patch: Fresh context, artifact + predicate + hunt list only; build history withheld; admit-or-reject, never repair (04-verifier.md).
8. Lenient holistic judgement
Exploit: Verifier skims, awards a pass on overall impression — rigor-looking setups without complete deduction survive.
Patch: Per-criterion citation required (a check with no citation doesn't count); graded 0–2 not binary; fail-closed so ambiguity → REJECT (04-verifier.md).
9. Circular satisfaction
Exploit: A step assumes something equivalent in strength to the goal — the subtlest near miss, and the one generic review never catches. Patch: Circularity is a named, mandatory section (C) in the verifier's hunt list, not a footnote.
10. Launder the answer from external sources
Exploit: Search for the solution, then present it as independent work; or conclude "this is unsolved" from a lookup and return that.
Patch: Retrieval scope + framing guard + external-query log retained for audit (brief-template.md, ledger EXTERNAL QUERIES).
11. Silent dropping of a blocked item
Exploit: Item 11 fails; the final summary covers 10 items and never mentions it. Count reads complete.
Patch: Blocked items marked [!] in the ledger and must be resolved or explicitly descoped with a logged scope change; "if an item cannot be completed, state it under its number — never fold it into a summary" (01-single-agent.md, ledger integrity rules).
12. Falsify the ledger's VERIFIED column
Exploit: Worker flips its own claims to VERIFIED without evidence — status theater with extra steps.
Patch: Evidence must be a session tool result / artifact hash / commit sha; only the verifier or harness flips the row; ledger checked before any completion claim (02-progress-ledger.md).
13. Halt on unanimity
Exploit: All workers agree → orchestrator returns. Agreement is weakest evidence exactly where it feels strongest — on the hardest problems.
Patch: Agreement is explicitly a diversity-failure signal, never a halt condition; audit content; strongest candidate wins on evidence, not on elegance or number of supporters (02-orchestrator.md).
Structural exploits considered and closed by design
| Class | Why it doesn't land |
|---|---|
| Plan / survey instead of artifact | Reporting contract rejects them by name |
| Reduction to an unproven equivalent | Non-counting list + circularity section |
| Bounded / sample verification | Non-counting list |
| Optimism without evidence | Evidence rule: claim without a pointer isn't reported |
| Endless run (no stop) | Harness budget: turns + cost + wall-clock |
| Duplicate work / re-doing merged work | Audit-first clause + falsified-hypotheses registry |
| Unbounded self-feed (tickets generated faster than closed) | Bounded self-feed + review cadence + parked backlog |
| Same fix repeated | Failed-strategies section + forced diagnosis before retry |
Round 2 — re-question
"How could an agent satisfy the letter of this brief without solving the problem?"
Against the patched kit, remaining answers are variations on the closed classes above (re-parameterized gate-loosening, re-worded scope narrowing). No new credible class emerged.
Standing rule: repeat this pass after any change to the task predicate, the DOES NOT COUNT list, or the verification gate. A patch to one can reopen a hole in another — exploit #1 was only reachable because "make tests green" had been added as a definition of done.