prompts.mikee.pro

QA 02 — Red-Team Pass

Method: the brief is handed to a fresh model instance with one question:

"How could an agent satisfy the letter of this brief without solving the problem?"

Below: every credible exploit found against the v1 templates, and the patch applied. Re-run this pass whenever the task predicate, non-counting list, or verification gate changes.


Round 1 — 13 exploits found against v1

1. Delete or weaken the tests to make them green

Exploit: npm test exits 0 because assertions were removed, tests .skipped, or coverage narrowed. The letter of "all tests pass" is satisfied. Patch: PreToolUse hook blocks test-file deletions and .skip/.only additions; Stop hook checks gate-config diffs; gate-loosening explicitly listed in DOES NOT COUNT (01-single-agent.md). Verifying the gate is itself un-gamed is part of the definition of done.

2. Satisfy the count contract by redefining items

Exploit: Produce a 7-item plan, then quietly broaden or narrow items 5–7 so less work counts as "done." Reports "7 of 7" truthfully against a moving target. Patch: Plan is frozen once produced; any change to an item's meaning must be logged as a SCOPE CHANGE under that item (01-single-agent.md, ledger SCOPE CHANGES).

3. Narrow the gate's scope instead of passing it

Exploit: Type-check passes because files were added to exclude, lint passes because a path was ignored, tests pass because the run was scoped to one passing suite. Patch: Definition of Done requires the full suite; touched-file lint derived from git diff; gate-config modifications require explicit sign-off (01-loop.md §5).

4. Match the harness's own stop signal on the agent's prose

Exploit: Agent ends its message with "Done." and a naive gate that greps output accepts it. Patch: Gates check exit codes and state, never text — including the agent's final message (01-loop.md §4).

5. Use the fallback clause as an escape hatch

Exploit: Return "the strongest derivation and its remaining gap" whenever the work gets hard — the fallback was at agent discretion. Patch: Fallback scoped to externally enforced budget exhaustion only, never agent discretion, clearly labeled incomplete (brief-template.md, 02-orchestrator.md).

6. Treat the effort floor as satisfaction of the return condition

Exploit: "8 hours have passed" becomes the reason to return, converting a floor into a timer. Patch: Explicit language everywhere: the floor revokes permission to quit early; it neither guarantees nor bounds runtime; only the artifact predicate satisfies the return condition (01-single-agent.md, 02-orchestrator.md, 01-loop.md §7).

7. Verifier collusion via shared context

Exploit: Verifier sees the build history and inherits the author's framing — "the remaining step is routine." Patch: Fresh context, artifact + predicate + hunt list only; build history withheld; admit-or-reject, never repair (04-verifier.md).

8. Lenient holistic judgement

Exploit: Verifier skims, awards a pass on overall impression — rigor-looking setups without complete deduction survive. Patch: Per-criterion citation required (a check with no citation doesn't count); graded 0–2 not binary; fail-closed so ambiguity → REJECT (04-verifier.md).

9. Circular satisfaction

Exploit: A step assumes something equivalent in strength to the goal — the subtlest near miss, and the one generic review never catches. Patch: Circularity is a named, mandatory section (C) in the verifier's hunt list, not a footnote.

10. Launder the answer from external sources

Exploit: Search for the solution, then present it as independent work; or conclude "this is unsolved" from a lookup and return that. Patch: Retrieval scope + framing guard + external-query log retained for audit (brief-template.md, ledger EXTERNAL QUERIES).

11. Silent dropping of a blocked item

Exploit: Item 11 fails; the final summary covers 10 items and never mentions it. Count reads complete. Patch: Blocked items marked [!] in the ledger and must be resolved or explicitly descoped with a logged scope change; "if an item cannot be completed, state it under its number — never fold it into a summary" (01-single-agent.md, ledger integrity rules).

12. Falsify the ledger's VERIFIED column

Exploit: Worker flips its own claims to VERIFIED without evidence — status theater with extra steps. Patch: Evidence must be a session tool result / artifact hash / commit sha; only the verifier or harness flips the row; ledger checked before any completion claim (02-progress-ledger.md).

13. Halt on unanimity

Exploit: All workers agree → orchestrator returns. Agreement is weakest evidence exactly where it feels strongest — on the hardest problems. Patch: Agreement is explicitly a diversity-failure signal, never a halt condition; audit content; strongest candidate wins on evidence, not on elegance or number of supporters (02-orchestrator.md).


Structural exploits considered and closed by design

Class Why it doesn't land
Plan / survey instead of artifact Reporting contract rejects them by name
Reduction to an unproven equivalent Non-counting list + circularity section
Bounded / sample verification Non-counting list
Optimism without evidence Evidence rule: claim without a pointer isn't reported
Endless run (no stop) Harness budget: turns + cost + wall-clock
Duplicate work / re-doing merged work Audit-first clause + falsified-hypotheses registry
Unbounded self-feed (tickets generated faster than closed) Bounded self-feed + review cadence + parked backlog
Same fix repeated Failed-strategies section + forced diagnosis before retry

Round 2 — re-question

"How could an agent satisfy the letter of this brief without solving the problem?"

Against the patched kit, remaining answers are variations on the closed classes above (re-parameterized gate-loosening, re-worded scope narrowing). No new credible class emerged.

Standing rule: repeat this pass after any change to the task predicate, the DOES NOT COUNT list, or the verification gate. A patch to one can reopen a hole in another — exploit #1 was only reachable because "make tests green" had been added as a definition of done.