Master Pseudo-Formal Brief Template
Copy, delete what doesn't apply, fill the rest. Blocks marked (parallel) are only needed when an orchestrator manages concurrent workers.
Order matters. Write the success predicate first. If you cannot write it in one sentence with explicit quantifiers and scope, the problem is not ready for a long-horizon run — decompose it or run a scoping session instead.
DEFINITIONS
<Every load-bearing term an adversarial reader could interpret two
ways. Include degenerate and boundary cases explicitly: the empty
input, the trivial solution, the duplicate, the disconnected case,
the zero-measurement. In empirical domains: units, populations,
inclusion criteria, measurement procedure.>
Definitions are loophole closure, not pedagogy. Each one should
pre-empt a specific way a lazy solution could argue either way.
TASK
<One statement of the success predicate: what must be true of the
returned artifact, with quantifiers and scope spelled out.>
<State it TWICE: once as the goal, once as the exact obligation —
and enumerate the assumptions the solution is NOT allowed to make.
Those enumerations are the special cases where partial results
are already known, so they block the most probable near misses.>
<If a solution plausibly exists:>
Assume for purposes of this task that a complete solution exists.
<If existence is genuinely uncertain — genuinely open questions,
ill-posed problems:>
Either a complete solution or a complete demonstration of
impossibility counts; nothing in between does.
<Never use "assume it exists" on a genuinely open question — the
run will fabricate rather than conclude impossibility.>
DOES NOT COUNT
Partial progress does not count unless it implies exactly the
resolution above. Specifically insufficient:
- results holding only for a narrowed scope or special case
- reductions to another unvalidated assumption or unproved statement
- verification over any bounded subset of cases
- artifacts with a requirement satisfied approximately where exact
satisfaction is specified
- candidate counterexamples without a complete certificate
- plans, surveys, status summaries, or explanations of difficulty
<THE REFUSAL LIST — the highest-leverage block. Imagine a capable
junior collaborator returning with each plausible partial result
and write down every one you would send back. That list, verbatim,
goes here. Predict the specific answer-shaped near misses THIS
problem invites and exclude each by name.>
ORCHESTRATION (parallel runs only)
Use concurrent agents aggressively and dynamically. Do not use
fixed assignments such as "N agents for strategy X". Heuristics:
- Begin with a genuinely diverse portfolio of substantially
different formulations: <list the known approach families for
this domain>.
- Do not tell most agents the currently favored approach.
Preserve independence in early rounds.
- Maintain an explicit registry of approach families, grouped by
the underlying IDEA rather than surface wording — paraphrase must
not be mistaken for diversity. Redirect agents away from crowded
families toward underexplored ones.
- Do not let one approach dominate because it yields elegant
reformulations. A route ending at a subproblem as hard as the
original goal is NOT progress unless it genuinely resolves it.
- When a route stalls at a goal-strength gap, mark it BLOCKED and
record why. Reassign only for a materially new mechanism,
invariant, or construction — never for renewed enthusiasm.
- Keep several incompatible routes alive across rounds;
cross-pollinate only after independent development has exposed
each route's real strengths and gaps.
- The root agent repeatedly synthesizes, challenges, redirects,
and launches new rounds. Do not stop after the first wave fails.
- Every spawn specifies: objective, output format, tool/source
guidance, and task boundaries. Missing any one produces
overlapping or gap-ridden coverage.
- Treat inter-agent agreement as a diversity-FAILURE signal, not
confirmation. Committees converge tightest on the hardest
problems. Audit content; never halt on unanimity alone.
VERIFICATION
Use adversarial reviewer agents with FRESH CONTEXT throughout —
a verifier that did not build the artifact cannot rationalize its
gaps. Every candidate is checked against this list:
<Enumerate the domain-specific ways a candidate can look right and
be wrong: known confounders, degenerate cases, circular arguments,
leakage paths, too-good-to-be-true signatures.
ALWAYS include the domain's version of circularity: satisfying
the goal by assuming something equivalent to it.>
- Require concrete artifacts: <lemmas, constructions, scripts,
datasets, measurements, counterexamples — domain-appropriate>.
Reject status reports, vague optimism, and claims that an
unresolved step is "routine".
- Grade each criterion 0-2 against the checklist; never binary.
- Structure the final artifact modularly so each part can be
verified in isolation, with premises and conclusion stated
locally.
- The verifier ADMITS or REJECTS. It never repairs the artifact —
if it repairs, checking collapses into self-repair.
REPORTING CONTRACT
- Every progress claim must trace to a tool result or artifact
from the current session. A claim you cannot point at evidence
for is not reported.
- Concrete artifacts only. Status reports and "on track" without
a pointer are rejected.
- If an item cannot be completed, state it explicitly under its
number — never silently drop it or fold it into a summary.
RETURN CONDITION
Return only when a candidate satisfies the TASK predicate and
survives the adversarial audit above. Do not return a reduction,
partial result, isolated missing step, best-effort summary, or
explanation of why the problem is difficult.
The return gate is checked by <verifier / harness>, not by your
own confidence.
<Fallback, scoped to external budget exhaustion ONLY — never at
the agent's discretion, or it becomes the escape hatch the rest
of the brief closed:>
If the externally enforced budget is exhausted first, return the
strongest rigorously verified derivation and its exact remaining
gap, clearly labeled as incomplete.
EFFORT
Spend at least <floor> before considering returning or giving up.
Do not return merely because current approaches fail; launch new
rounds and search for fresh formulations.
<An effort floor is a PERMISSION REVOCATION, not a schedule. It
removes permission to quit early; it neither guarantees nor
bounds runtime. Never let elapsed time satisfy the return
condition — only the artifact predicate does.>
CONTAMINATION
External search may be used only for <background, standard named
results, documented APIs>. Do not search for a solution to this
exact problem or its benchmark.
<If solvability framing is used:>
Do not conclude from external sources that the problem is unsolved,
and do not answer that it is open.
<Retain a log of external queries so independence is checkable.>
Filling notes
- The refusal-list method is the fastest route to a strong DOES NOT COUNT block.
- Solvability framing is a scalpel. Engineering problems and well-evidenced conjectures → use it. Genuinely open questions → use the two-sided form.
- The fallback clause must be scoped to external budget exhaustion, or it reopens the door the brief closed.
- Effort floors are permissions, not schedules. Enforce real time and cost limits in the harness; a prompt cannot bound spend.
- Keep hard constraints out of the brief. Budgets, tool permissions, sandbox boundaries → runtime. Mention them only for planning.
- Per-spawn specs: objective, output format, tool guidance, boundaries — all four, every spawn.
Pre-launch rubric
See qa/pre-launch-rubric.md. Score 0/1/2 per dimension; fix every 0 and 1 before launch.
Red-team before launch
See qa/red-team.md. Give the finished brief to a fresh model instance and ask one question:
"How could an agent satisfy the letter of this brief without solving the problem?"
Patch every credible answer. Repeat until the answers stop being credible.