Harness 02 — Verified Progress Ledger
Why this exists: externally maintained verified-progress state, re-injected each round, is what rescued tasks that prompt-only and completion-gated setups failed entirely in controlled comparison. Prompt-only progress claims drift; a ledger re-read from disk does not.
The rule that makes it trustworthy:
An entry enters
VERIFIEDonly when it points at a tool result or artifact from the current session. A claim you cannot point at is not progress.
Only the verifier (or the harness) flips a row to VERIFIED. The worker proposes; it does not grade its own work.
# PROGRESS LEDGER — <run id>
_Last re-injected: <timestamp> | Budget remaining: <turns> turns, $<n>, <h>h_
## TASK PREDICATE
<verbatim success predicate — re-stated every round, because a
budget stated once loses force as context grows>
## COUNT
_plan: 12 items — done: 7 — blocked: 2 — remaining: 3_
## PLAN
- [x] 1. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 2. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 3. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 4. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 5. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 6. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [x] 7. <item> → evidence: `<command>` exit 0 @ <file/commit>
- [ ] 8. <item> — in progress, next action: <one specific tool action>
- [ ] 9. <item>
- [ ] 10. <item>
- [!] 11. <item> — BLOCKED: <exact missing step> / <what was tried>
- [!] 12. <item> — BLOCKED: <exact missing step> / <what was tried>
## VERIFIED RESULTS
| # | Claim | Evidence (session tool result / artifact) | Verified by |
|---|-------|--------------------------------------------|-------------|
| V1 | <claim> | `<command output / file hash / commit sha>` | verifier |
| V2 | <claim> | `<command output / file hash / commit sha>` | verifier |
## APPROACH REGISTRY (grouped by IDEA, not wording)
| Family | Status | Owner | Notes |
|--------|--------|-------|-------|
| <approach A> | active | agent-3 | <real strengths / gaps found> |
| <approach B> | active | agent-7 | |
| <approach C> | crowded → redirected | — | 4 workers converged; 2 reassigned |
| <approach D> | BLOCKED | — | stalls at <goal-strength gap>; reopen only for materially new mechanism |
## FALSIFIED HYPOTHESES (do not re-explore)
- <hypothesis> — falsified by <evidence>, round <n>
- <hypothesis> — falsified by <evidence>, round <n>
## FAILED STRATEGIES (do not repeat)
- <attempt> → <exact error / result>. Do not retry this.
- <attempt> → <exact error / result>. Do not retry this.
## SCOPE CHANGES
- item 5: redefined from <old> to <new> — reason, round <n>
<A scope change is always logged. Silent redefinition is how a
count contract gets satisfied without the work being done.>
## DECISIONS LOG (for human review at the end)
- <decision> — chose <A> over <B> because <reason>; flag: <needs review?>
- <judgment call> — call made: <what>; alternatives: <both options>
## EXTERNAL QUERIES (contamination guard)
- <query> — purpose: background / named result ✓
- <query> — purpose: <flag if it looks like searching for the answer>
## BACKLOG (self-fed work, untriaged unless blocking)
- [ ] <ticket> — linked to epic, actioned: yes/no
- [ ] <ticket> — parked
## NEXT ROUND SHOULD
<one or two concrete next actions — a tool action, not a high-level intention>
Why each section is here
| Section | Failure it prevents |
|---|---|
| Evidence column | Fabricated status reports — claims must trace to a session tool result |
[!] blocked items |
Blocked work silently absorbed into "complete" |
| Approach registry | Paraphrase mistaken for diversity; one family dominating |
| Falsified hypotheses | Parallel workers re-exploring dead ends |
| Failed strategies | Same fix attempted repeatedly |
| Scope changes | Count contract satisfied by quietly redefining the item |
| Decisions log | Interruption-by-question — autonomy granted, judgment calls batched for review |
| External queries | Result laundering |
| Next round | Vague "continue" that restarts from step 1 after compaction |
Injection points
- Start of every round — orchestrator re-reads before spawning
- After every compaction / context reset — before the next action
- Session-start ritual — read ledger + git log → smoke-test → pick exactly one unfinished item
- Before any completion claim — verifier checks the claim against the ledger, not against the agent's memory
Integrity rules
- Worker proposes
VERIFIED; verifier flips it, citing the artifact - Deleting or editing tests to flip a row is a gate-loosening action → blocked by PreToolUse hook
- The ledger is append-mostly; removals require a logged reason
DONErequires: every plan item[x], every[!]resolved or explicitly descoped with a logged scope change, and the definition-of-done checks green