PAIN-POINTS:
P1 Unsettled contradiction policy
  evidence: st-40 cand — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:21; st-40 base — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:57; st-30 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-TABLE.md:5; st-30 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-TABLE.md:6
  root cause: Frozen plans do not classify a defect as fatal, uniquely resolvable, an ordered fallback, or local to one action. The executor therefore makes a routing-policy decision, and the judge families reward opposite choices.
  owner layer: planner-side
  why prompt text cannot fix it: Executor wording cannot settle an operator policy that remains disputed by the evaluators themselves.

P2 Untyped plan-to-code identity
  evidence: st-24 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:18; tv-13 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:21; tv-23 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:31; st-57 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:14
  root cause: Exact routines, constants, input constructions, thresholds, and fallback permissions exist only in free Markdown. The same executor interprets that prose, implements it, and audits its own interpretation, so “scientifically nearby” substitutions repeatedly pass its self-check.
  owner layer: schema
  why prompt text cannot fix it: More examples still cannot mechanically establish that implemented code denotes the same routine, population, constant, and law as the frozen plan.

P3 Claim provenance is not first-class
  evidence: tv-13 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:23; st-30 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:10; st-19 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:11; complaints-ledger row b10-ctl1-V1
  root cause: The runner records file paths, shapes, and declared grain, but not claim IDs or field-level bindings; result prose remains an unconstrained second computation surface. Consequently a quotient, range, difference, or diagnostic can be true, false, or unreproducible without any artifact-level distinction.
  owner layer: schema
  why prompt text cannot fix it: K50a was explicit and still breached because compliance depends on producing and consuming a typed field, not remembering a sentence while writing prose.

P4 Ephemeral preflight proof
  evidence: st-30 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:47; st-57 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:51; st-40 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:58
  root cause: The executor is asked to test proposed repairs on synthetic shapes, yet the only authorized scratch channel must be deleted before delivery and has no persistent synthetic-evidence type. A failure decision can therefore depend on evidence the reviewer is structurally unable to inspect.
  owner layer: runner
  why prompt text cannot fix it: Exact obedience to the scratch-deletion rule destroys the proof; an artifact channel must exist before wording can require its use.

P5 Soft data boundary
  evidence: st-24 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:15; st-24 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:17; complaints-ledger row b10-p66-V1
  root cause: Governed data is physically readable by the main executor and coder; the runner governs only commands voluntarily routed through it. A profiling or scratch invocation therefore has the same filesystem capability as a governed action.
  owner layer: runner
  why prompt text cannot fix it: The boundary was already stated explicitly when both breaches occurred; a behavioral prohibition cannot replace capability isolation.

P6 Compute has no admission control
  evidence: st-24 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:19; tv-64 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:32; st-30 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:8; st-40 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:56
  root cause: The plan budget and 6,541-word implementation catalogue are advisory prose; no pre-dispatch artifact compares actual fit calls, diagnostics, materializations, or algorithms with the declared inventory. Optional audits and hand-rolled machinery can therefore enter source before their cost is exposed by timeout.
  owner layer: skill
  why prompt text cannot fix it: The existing prompt and skill already forbid and price these operations; the missing operation is a plan-specialized source/cost check.

P7 Nontransactional action packages
  evidence: st-19 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:5; st-19 — evaluation/meta-harness/outputs/2026-08-29/22-07-c-judge-c1-executor/ROUND-FINDINGS.md:10; tv-23 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:32; tv-13 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:57
  root cause: The runner records attempts, but the stop gate validates terminal shape and successful pointers rather than a closed action transaction: executed source snapshot, live attempt, complete manifest, expected files, and matching action report. Honest failure terminals are valid, but nothing prevents them from citing stale or absent products as empirical support.
  owner layer: stop-gate
  why prompt text cannot fix it: Only rerunning or validating the final package state can detect a surviving ShapeError, overwritten evidence, or empty live manifest.

P8 Self-authored reconciliation
  evidence: st-57 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:18; st-57 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:20; tv-23 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:61; tv-23 — evaluation/meta-harness/outputs/2026-08-29/23-24-c-judge-c2-executor/ROUND-FINDINGS.md:64
  root cause: One model writes source, audits it, writes each action report, and manually restates those reports in the terminal after a large instruction load. No independent semantic reconciliation occurs until the downstream reviewer, so false CI statements, omitted deviations, and invalid posterior uses survive executor delivery.
  owner layer: reviewer-side
  why prompt text cannot fix it: Repeating an honesty duty does not make the author of the disputed statement an independent verifier.

ARCHITECTURE-OPTIONS:
O1 Typed execution contract
  mechanism: Add a planner-owned execution-contract.yaml validated by a new schema beside plan.md. Each Axx records exact input construction, routine/API and parameters, constants, gate condition and rationale grains, ordered/forbidden fallbacks, output fields and grain, fit/pass/repetition inventory, and the operator-selected contradiction policy; render the human action block from this object so two versions cannot drift.
  covers: P1, P2
  cost: Roughly 700–1,100 LOC plus migration tooling; no new doctrine text, although the executor reads a compact structured artifact that should replace substantial prompt prose.
  risk: A schema can create false confidence while leaving scientific validity unchecked; dual-authoring would introduce drift, so plan.md must be rendered from or directly reference the structured object.
  proof: Replay st-30, st-40, st-57, and tv-64 after operator labels each disputed case; success means deterministic route agreement, zero unlisted substitutions, and no Codex/Opus split caused by missing policy.

O2 Implementation broker and independent source auditor
  mechanism: Replace the long choose-implementation catalogue with a short skill backed by a queryable capability/benchmark registry. It reads execution-contract.yaml and writes implementation-choice.json containing compatible APIs, forbidden near-neighbours, expected calls/materializations, and estimated cost; a mandatory source-auditor sub-agent independently maps source call sites to that record before dispatch.
  covers: P2, P6
  cost: Roughly 500–800 LOC, a 300–500-word skill and sub-agent brief; it should remove about 6,000 words from the executor’s visible instruction load rather than add to it.
  risk: The registry can become stale, static call accounting misses dynamic dispatch, and an auditor model can reproduce the executor’s semantic mistake; legitimate custom implementations need an explicit planner authorization path.
  proof: Seed the st-24 hand likelihood, st-57 weighted-score solver, tv-13 thresholds, tv-23 window rule, and extra-fit variants; require pre-dispatch detection above 95%, zero false authorization of substitutions, and predicted versus governed cost within a preregistered tolerance.

O3 Field-level claim ledger and generated terminal
  mechanism: Require every action to persist an aggregate claims table with stable claim IDs, source file/field, value, unit, attempt, and provenance. Action prose uses claim tokens, and a result assembler builds numeric tables, posterior inputs, deviation union, and terminal reconciliation from live action records; arbitrary prose may not introduce empirical numerals.
  covers: P3, P8
  cost: Roughly 600–900 LOC and under 100 words telling the executor to invoke the assembler; no additional conduct doctrine.
  risk: Dates, IDs, plan constants, and mathematical examples need separate token types; rigid generation may reduce readability, and an omitted claim still needs reviewer detection.
  proof: Run the bank over c1, c2, and B-R10: zero unmatched empirical numerals, 100% field-open success, and seeded false CI/deviation/posterior inconsistencies rejected before terminal creation.

O4 Capability-scoped runner with persistent synthetic mode
  mechanism: Keep governed data unmounted or unreadable in the main executor/coder context and expose it only inside the runner child’s mount namespace. Add a synthetic-preflight mode with no governed-data capability that persists generator source, seed, synthetic inputs, aggregate outputs, and provenance under workspace/preflight/ for reviewer inspection.
  covers: P4, P5
  cost: Roughly 500–900 LOC plus container/mount integration; essentially no executor instruction growth beyond naming the two runner modes.
  risk: Debugging becomes harder, packages may assume globally visible data paths, and synthetic artifacts could be mistaken for empirical evidence unless provenance is prominent and schema-enforced.
  proof: Replay the c1-st24 and b10-p66 out-of-runner commands and require OS-level denial, while st-30/st-40/st-57 synthetic justifications remain openable and reproduce exactly.

O5 Narrow transactional package gate
  mechanism: Archive an immutable source-tree snapshot and digest per runner attempt, emit package-state.json, and extend result validation to check current-attempt source identity, exact live manifest membership, nonempty evidence for success, action-result/terminal status equality, and claim pointers that do not reach archived or missing files. The operator’s broader mechanical conduct checker was abandoned; this does not resurrect it—it checks only closed transaction invariants already represented by typed artifacts, not chained commands, out-of-runner conduct, or scientific rules.
  covers: P7, P3, P8
  cost: Roughly 250–450 LOC; no instruction text added to the executor.
  risk: Multi-entry sources and legitimate failure packages need careful modeling; expanding it into semantic source inspection would recreate the rejected rule engine.
  proof: Mutation-test the empty manifest, changed-after-run source, stale archived pointer, overwritten segment file, and terminal/action status mismatch; all must fail, while contradiction-only and honest implementation-failure terminals still pass.

O6 Reviewer-owned execution reconciler
  mechanism: Add a mandatory d3_reviewer execution-reconciler sub-agent that reads the typed contract, executed source snapshot, live evidence claims, action reports, and terminal, then writes a structured reconciliation under committee/. It independently rederives decisive regions and enumerates every executed deviation; d3_reviewer cannot accept while reconciliation has an unresolved mismatch.
  covers: P8, P2
  cost: Roughly 150–250 LOC plus a 300–500-word reviewer-only sub-agent brief; zero executor instruction growth.
  risk: Detection occurs after compute is spent, reviewer latency rises, and another model can still miss a scientific error; it is a final semantic backstop, not execution control.
  proof: Inject the st-57 omitted 0/199 convergence fact, tv-23 false CI statement, invalid posterior, and c1-st24 omitted substitution; require 100% rejection before Atom production.

RANKED:
O1 — It settles the missing policy and creates the typed foundation every later checker, skill, runner, and reviewer needs.
O3 — It directly removes the highest-frequency evidence/reporting surface with deterministic field-level provenance instead of another behavioral rule.
O2 — It attacks substitution and budget blow-ups while materially shrinking, rather than enlarging, the executor’s instruction context.

POLICY-QUESTIONS:
Q1 Should a literal inconsistency with a uniquely named conventional reading be refused or executed with disclosure? On st-40, Codex penalized cand for executing canonical Webb (C1 ROUND-FINDINGS.md:21) while Opus penalized base for refusing it (C1 ROUND-FINDINGS.md:57); this is an unresolved policy, not unclear executor wording.
Q2 Must independent actions and checks execute when another input or branch is blocked? Opus faulted tv-64 for suppressing three independent checks (C1 ROUND-FINDINGS.md:63-64) and st-57 for running zero dispatches despite executable core work (C2 ROUND-FINDINGS.md:50-53); the planner contract must define blocker scope.
Q3 Is the policy “every derived empirical number must be persisted,” or may typed one-line derivations over cited aggregate fields appear? The b8-ctl2 and b9-ctl1 ledger rows record repeated Codex/Opus severity splits despite identical arithmetic; the operator must choose before claim-schema and evaluator behavior can align.

CHALLENGE:
The campaign did not reach its preregistered ten rounds or held-out evaluation, while an identical-prompt A/A pair still scored 2–0 (A/A ROUND-TABLE.md:5-12); two four-problem rounds cannot support a stable improvement estimate.
The treatment changed 5,166 words to 3,806, but the executor still receives roughly 12,800 total words dominated by the 6,541-word implementation skill; this tested one document, not the instruction architecture.
The judge brief makes efficiency conditional on correctness and treats substitution as fatal, so executing an ambiguous plan creates many more losing surfaces than refusing it; the comparison partly measures risk posture.
Overall dimension weights are unspecified, and the same st-30/st-57 pairs yield Codex-base versus Opus-cand (C2 ROUND-TABLE.md:5-8); “prompt quality” is confounded with judge-family policy.