prompts.mikee.pro

QA 01 — Pre-Launch Rubric (scored)

Score each dimension 0 (absent), 1 (present but gameable), 2 (adversary-proof). Fix every 0 and 1 before launch. Expensive runs deserve a passing brief.

Scores below are for the kit as shipped in ../.


# Dimension 2 means Score Why it scores 2
1 Success predicate Adversarial reader decides unambiguously; quantifiers and scope explicit 2 Stated twice (goal + exact obligation), forbidden assumptions enumerated by name
2 Definitions Every load-bearing term defined; degenerate cases settled 2 Definitions block specific loopholes; verified = exit code read, not stdout text
3 Non-counting outcomes This problem's plausible near misses excluded by name 2 Refusal-list method built in; gate-loosening, scope-narrowing, silent drops all named
4 Auditor checklist Enumerated, domain-specific failure modes incl. circularity analogue 2 04-verifier.md sections C + D; per-criterion citation required
5 Persistence–verification pairing Every persistence instruction has a matching verification gate 2 Autonomy & IF BLOCKED both carry an explicit <PAIRED GATE>; harness Stop hook is the enforcement
6 Return condition Predicate over the artifact; fallback scoped to external budget exhaustion only 2 Return gated on audit; fallback only on harness budget exhaustion, never agent discretion
7 Diversity policy (parallel) Early independence, idea-keyed registry, blocked-route rules, late cross-pollination 2 All four present in 02-orchestrator.md, incl. anti-elegance rule and unanimity-is-not-confirmation
8 Reporting contract Concrete artifacts required; claims trace to session evidence 2 Evidence rule in prompt + VERIFIED/evidence columns in ledger
9 Contamination guards Retrieval scope stated wherever result independence matters 2 Scope + framing guard + external-query log
10 Harness separation No hard constraint lives only in the prompt 2 Budgets, permissions, stop gates, ledger injection all in ../harness/

Total: 20/20.


Supplementary checklist (from research beyond the original 10)


Red-team status

Full pass recorded in 02-red-team.md: 13 credible exploits found against v1, 13 patched.

The standing instruction on any brief you adapt from this kit:

Give the finished brief to a fresh model instance and ask: "How could an agent satisfy the letter of this brief without solving the problem?" Patch every credible answer. Repeat until the answers stop being credible.

Re-run this whenever you change the task predicate, the non-counting list, or the verification gate — a patch to one can reopen a hole in another.


Launch decision

Check Status
All 10 rubric dimensions at 2 ✅
Supplementary checklist complete ✅
Red-team exploits patched ✅ 13/13
Harness enforcement in place (stop gate, budget, ledger) ⚠️ must be configured before launch
Fresh verifier context confirmed ⚠️ must be confirmed before launch

Conditional go. The prompt layer is ready; the run is not launch-safe until the Stop hook, budget, and ledger injection are actually wired up in the runtime. A brief with a gate that exists only on paper is the failure mode this whole methodology exists to prevent.