QA 01 — Pre-Launch Rubric (scored)
Score each dimension 0 (absent), 1 (present but gameable), 2 (adversary-proof). Fix every 0 and 1 before launch. Expensive runs deserve a passing brief.
Scores below are for the kit as shipped in ../.
| # | Dimension | 2 means | Score | Why it scores 2 |
|---|---|---|---|---|
| 1 | Success predicate | Adversarial reader decides unambiguously; quantifiers and scope explicit | 2 | Stated twice (goal + exact obligation), forbidden assumptions enumerated by name |
| 2 | Definitions | Every load-bearing term defined; degenerate cases settled | 2 | Definitions block specific loopholes; verified = exit code read, not stdout text |
| 3 | Non-counting outcomes | This problem's plausible near misses excluded by name | 2 | Refusal-list method built in; gate-loosening, scope-narrowing, silent drops all named |
| 4 | Auditor checklist | Enumerated, domain-specific failure modes incl. circularity analogue | 2 | 04-verifier.md sections C + D; per-criterion citation required |
| 5 | Persistence–verification pairing | Every persistence instruction has a matching verification gate | 2 | Autonomy & IF BLOCKED both carry an explicit <PAIRED GATE>; harness Stop hook is the enforcement |
| 6 | Return condition | Predicate over the artifact; fallback scoped to external budget exhaustion only | 2 | Return gated on audit; fallback only on harness budget exhaustion, never agent discretion |
| 7 | Diversity policy (parallel) | Early independence, idea-keyed registry, blocked-route rules, late cross-pollination | 2 | All four present in 02-orchestrator.md, incl. anti-elegance rule and unanimity-is-not-confirmation |
| 8 | Reporting contract | Concrete artifacts required; claims trace to session evidence | 2 | Evidence rule in prompt + VERIFIED/evidence columns in ledger |
| 9 | Contamination guards | Retrieval scope stated wherever result independence matters | 2 | Scope + framing guard + external-query log |
| 10 | Harness separation | No hard constraint lives only in the prompt | 2 | Budgets, permissions, stop gates, ledger injection all in ../harness/ |
Total: 20/20.
Supplementary checklist (from research beyond the original 10)
- Stop condition writable as a check (exit code, test count, state) —
01-loop.md§3 - Verifier runs in fresh context, separate from builder —
04-verifier.md - External ledger re-injected each round, not carried in prompt —
02-progress-ledger.md - Hard budgets enforced by runtime, not prompt —
01-loop.md§7 - Ambient "done" text disarmed as a stop signal —
01-loop.md§4 - Count contract ("7 of 12 done") for multi-part asks —
01-single-agent.md - Unresolved constraints restated with mapped next action (mitigates measured late-stage pressure) — ledger
NEXT ROUND SHOULD+ plan items carry next action - Every spawn carries objective + output format + tool guidance + boundaries —
03-worker-spec.md - Gate integrity — PreToolUse blocks weakening tests/configs
- Inter-agent agreement never used as a halt signal — orchestrator + verifier
Red-team status
Full pass recorded in 02-red-team.md: 13 credible exploits found against v1, 13 patched.
The standing instruction on any brief you adapt from this kit:
Give the finished brief to a fresh model instance and ask: "How could an agent satisfy the letter of this brief without solving the problem?" Patch every credible answer. Repeat until the answers stop being credible.
Re-run this whenever you change the task predicate, the non-counting list, or the verification gate — a patch to one can reopen a hole in another.
Launch decision
| Check | Status |
|---|---|
| All 10 rubric dimensions at 2 | ✅ |
| Supplementary checklist complete | ✅ |
| Red-team exploits patched | ✅ 13/13 |
| Harness enforcement in place (stop gate, budget, ledger) | ⚠️ must be configured before launch |
| Fresh verifier context confirmed | ⚠️ must be confirmed before launch |
Conditional go. The prompt layer is ready; the run is not launch-safe until the Stop hook, budget, and ledger injection are actually wired up in the runtime. A brief with a gate that exists only on paper is the failure mode this whole methodology exists to prevent.