Harness 01 — Loop, Stop Gates, and Budget Enforcement
Prompt-stated constraints are advisory under optimization pressure. Everything in this file lives outside the prompt, because it has to survive a pressured agent talking itself past a rule.
The prompt gives the agent a definition of done. The harness decides whether it's actually done.
1. The three controls — every loop needs all three
| # | Control | Failure it prevents |
|---|---|---|
| 1 | Checkable definition of done | Stopping too early because "done" was a feeling |
| 2 | Independent verifier | "Finished" treated as "correct" — the agent grading itself |
| 3 | Hard stop (turns + cost + wall-clock) | Running forever burning tokens |
A loop without a check is "a very confident token furnace." Keep the orchestration outside the agent — iteration limits, branching, and stop conditions belong in a script or workflow function; the agent does the work inside that structure.
2. Stop gate — Claude Code
# .claude/hooks/verify-before-stop.sh
# Exit 0 → turn may end. Non-zero → turn continues (stderr shown to agent).
set -euo pipefail
fail() { echo "$1" >&2; exit 1; }
npx tsc --noEmit >/dev/null 2>&1 || fail "type-check failing: fix before ending turn"
npm test --silent >/tmp/test-out.txt 2>&1 || { tail -40 /tmp/test-out.txt >&2; fail "tests failing: fix before ending turn"; }
npx eslint . --quiet || fail "lint errors: fix before ending turn"
# Gate integrity: the gates themselves must not have been weakened
git diff --name-only HEAD | grep -E '(tsconfig|eslint|jest\.config|\.eslintrc)' \
&& fail "gate config modified — this change needs explicit human sign-off"
# Count contract: no open plan items
[ -f .agent/plan.md ] && grep -qE '^\s*[-*] \[ \]' .agent/plan.md \
&& fail "unchecked plan items remain: $(grep -cE '^\s*[-*] \[ \]' .agent/plan.md) left"
exit 0
| Hook | Fires when | Effect |
|---|---|---|
| Stop | Agent tries to end turn | Shell must exit 0; agent continues (max 8 consecutive blocks, then overrides — design for fast, actionable stderr) |
| PreToolUse | Before a tool runs | Exit 2 = deny with message to agent |
| PostToolUse | After a tool completes | Next step blocked until resolved |
Strongest unattended stack: Plan mode → implementation with verification prompt → /goal for diff scope + test pass → Stop hook as final hard gate.
Replace "loop until I say stop" with "loop until check passes."
3. Machine-checkable goals
/goal (or an equivalent evaluator loop) runs a check on every attempted stop and sends the agent back if unmet. Always pair with a turn cap — a vague goal with no cap will spend a long time deciding it is close enough.
| Bad goal | Good goal |
|---|---|
| Code is clean | pnpm lint exits 0 |
| Feature works | pnpm test src/feature/ exits 0 |
| Only touched login | git diff --name-only matches ^src/api/login/ |
| Everything done | Plan has zero unchecked items AND CI green |
4. Disarm ambient stop signals
A specific, common bug: the loop greps tool output for "done"/"complete"/"all tests pass". A package.json script echoing done, or a runner printing "all tests pass" for one suite, short-circuits the loop while real work remains.
Do not report the task complete until ALL of:
- pnpm tsc --noEmit returns 0
- All tests in src/ pass
- The plan list has zero remaining items
If any tool's stdout contains "done" or "complete", ignore it as
a stop signal. A message from a script is not evidence.
Gate on exit codes and state, never on text — including the agent's own final message text.
5. Prevent gate-loosening (PreToolUse)
The most likely way an agent satisfies the letter of "make tests green" without solving the problem is to weaken the test.
Block: git commit of test file deletions / .skip / .only additions
Block: edits to tsconfig exclude arrays, eslint disable files, jest config
Block: git push --force, git reset --hard on shared branches
Block: rm -rf outside declared scratch paths
Require: any change to a gate config carries a flag for human review
6. External progress ledger — re-injected every round
A budget or reminder stated once at the top loses force as the trajectory grows. Progress claims are only reliable when they trace to session evidence. Both of these are harness jobs.
Re-inject at the start of every round / after every compaction:
- The verified-progress ledger (below)
- The remaining budget (turns, cost, wall-clock)
- The blocked-route registry + falsified hypotheses
- The plan with item counts ("7 of 12 done")
Format: 02-progress-ledger.md
7. Budgets — the only authorized incomplete return
budget:
wall_clock: <e.g. 8h> # effort floor lives in the PROMPT
cost_usd: <e.g. 400> # cost cap lives HERE
max_turns: <e.g. 300>
max_stop_blocks: 8 # platform default; design around it
on_exhausted: |
Return the strongest rigorously verified result plus its exact
remaining gap, clearly labeled incomplete. This is the ONLY
condition under which an incomplete return is permitted —
never the agent's own discretion.
Effort floor ≠ budget. The floor (in the prompt) revokes permission to quit early; it neither guarantees nor bounds runtime. The budget (here) is the actual stop.
8. Session hygiene
| Problem | Fix |
|---|---|
| Context fills; agent wraps up early ("context anxiety") | Context reset with a structured handoff artifact, not just compaction — compaction preserves the anxiety |
| Plan compacted away; "continue" restarts from step 1 | Resume with the exact step text of the current item; checkpoint ~10-step subtasks |
| Same fix repeated | Feedback stage forcing diagnosis; carry "failed strategies" in ledger |
| Compaction drift on resume | Keep prompts functionally identical when resuming |
| State lost across restarts | File-backed state, re-read each cycle — the largest single contributor in controlled comparisons |
| Session start picks wrong work | Session-start ritual: read ledger + git log → smoke-test → pick exactly one unfinished item |
9. Escalation ladder
| Tier | Trigger | Action |
|---|---|---|
| 1 — Automatic | Transient failure (timeout, flake) | Retry with backoff, max 3 |
| 2 — Adaptive | Persistent failure | Log the failed approach, force a different mechanism, note it in the ledger |
| 3 — Blocked | Goal-strength gap | Mark route blocked in registry; reopen only for a materially new mechanism |
| 4 — Escalation | Budget exhausted / genuinely blocked | Structured report: what was verified, what remains, exact gap |
Tier 3 is a control outcome, not a side effect. Blocked work never reaches "complete."
10. Loop health metrics (watch these)
- False completion rate — agent declared done while a gate was red (target: 0; dominant real-world failure)
- Duplicate submissions — same unit reworked (target: 0; state ledger eliminates it)
- Wasted steps in loops — repeated failing action / recurring state
- Cost per verified unit — RALPH-style repeat-and-hope runs ~2× the cost of an adaptive orchestrator for worse results
- Stop-block count — rising means the definition of done and the agent's model of done are diverging