Long-Horizon Prompt Kit
Ready-to-deploy prompts and harness config for runs that must keep working until all work is done — built from deep research into pseudo-formal task briefs, persistence/verification pairing, and what has to live outside the prompt to survive a pressured agent.
Governing rule: Never add a persistence instruction without a matching verification gate.
Full research report: ../long-horizon-prompting-research.md
Contents
The brief
| File | Use it for |
|---|---|
brief-template.md |
Master pseudo-formal brief — start here. 10 blocks, filling notes |
Prompts
| File | Use it for |
|---|---|
prompts/01-single-agent.md |
One agent on a coding/build task — autonomy + count contract + paired gates |
prompts/02-orchestrator.md |
Root agent managing parallel workers on an open-ended problem |
prompts/03-worker-spec.md |
The four-part spawn spec — objective, format, tools, boundaries |
prompts/04-verifier.md |
Fresh-context adversarial reviewer — admit/reject, never repair |
Harness (enforcement — not optional)
| File | Use it for |
|---|---|
harness/01-loop.md |
Stop hooks, /goal, gate integrity, budget, session hygiene, escalation |
harness/02-progress-ledger.md |
Verified-progress ledger re-injected every round |
QA
| File | Use it for |
|---|---|
qa/pre-launch-rubric.md |
10-dimension scorecard — 20/20, conditional go |
qa/red-team.md |
"How could an agent satisfy the letter without solving it?" — 13 exploits, 13 patched |
Launch sequence
1. Write the success predicate. One sentence, explicit scope.
└─ Can't write it? Don't launch. Decompose instead.
2. Fill the brief from brief-template.md
└─ Refusal list first: what would you send a junior colleague
back for? That list is your DOES NOT COUNT block.
3. Red-team it → qa/red-team.md method
└─ Fresh model, one question, patch every credible answer.
4. Score it → qa/pre-launch-rubric.md
└─ Fix every 0 and 1.
5. Wire the harness → harness/
└─ Stop hook · /goal · budget · ledger injection
└─ ⚠️ A gate that exists only on paper is not a gate.
6. Launch.
7. Verify → admit → only then return.
The four load-bearing pieces
If you take nothing else:
- Success predicate — one sentence, quantified, with forbidden assumptions enumerated. Everything else hangs off it.
- Non-counting outcomes — the refusal list. Under persistence pressure an agent returns an answer-shaped near miss; each named exclusion closes an escape hatch. Highest-leverage block in the brief.
- Return condition as an artifact predicate — "survives adversarial audit," never "I'm confident" or "the floor elapsed." Fallback only on externally enforced budget exhaustion.
- Harness-enforced gate — the prompt asks; the runtime decides. Budgets, permissions, stop conditions, and the ledger live outside because a prompt-stated constraint is advisory under pressure.
Non-negotiable pairings
| If you add… | You must also add… |
|---|---|
| "Keep going until…" | A machine-checkable definition of done |
| An effort floor | A return condition that elapsed time does not satisfy |
| Parallel workers | An idea-keyed approach registry + early independence |
| A verifier | Fresh context + enumerated failure modes + read-only admission |
| A fallback ("strongest result") | Budget-exhaustion-only scoping |
| A count contract | Frozen plan + logged scope changes |
| Progress claims | Evidence pointers from the current session |
| External search | Retrieval scope + query log |
Removing the right column while keeping the left is precisely the configuration that burns hours producing confident non-solutions.