Skip to content

Verdict

REVISE. The concept is accepted in principle and the governance work is sound. The code is not approvable and the gating-neutrality proof does not establish what it claims.

B1 — BLOCKING: the proposal cannot write a journal record

ops/tool-proposals/regime_snapshot.py.proposed opens the journal read-only and then writes to it:

-            with open(JOURNAL, "a") as f:     # deployed
+            with open(JOURNAL) as f:          # proposal
                 f.write(json.dumps(record) + "\n")

f.write() on a read-mode handle raises io.UnsupportedOperation: not writable. This is unrelated to provenance or taxonomy: it is an unrelated edit landing on the single line that commits a regime record to disk.

Reproduced against real inputs, append enabled, as production runs it:

path deployed proposal
healthy, --asof 2026-09-24 exit 0, record appended exit 1, traceback, nothing written
fail-closed, --asof 2026-09-27 exit 2 exit 1, traceback, nothing written

Applied to production this throws at 10:31 with no regime record written and an exit code the round has no branch for — neither the 0 that means "labelled" nor the 2 that means "fail-closed". Verified sha256 that the reviewed file is byte-identical to the one in /tmp/shadow on hermes, so this is the code that was proven.

B2 — BLOCKING: the proof could not have failed

/tmp/shadow/options-system-wiki/96-journal/ is empty — there is no regime.jsonl. Trace the code with that fact: os.path.exists(JOURNAL) is false, prior stays empty, the any(...) idempotency test is false, so control reaches the defective line on every run. It would raise FileNotFoundError there even before the mode error. The only way that sandbox yields the reported "exit codes 2 and 0" is --no-append — the one flag that skips the defective line.

So the proof compared record contents and never exercised the write path. It was structurally incapable of detecting this class of defect, which makes "gating neutrality proven by shadow run" a stronger claim than the evidence supports.

Required change to the method, not just the code: re-run the proof with append ENABLED, against a sandbox seeded with a copy of the real regime.jsonl, and diff the appended records — not stdout. A proof that runs in a mode production never uses is not a proof of production. State the exact command line used, including flags, in the record.

B3 — DESIGN: the taxonomy mislabels vendor failures as ours

Every har_forecast failure is hardcoded error_class: "internal". That script's failure modes are mixed:

har_forecast.py failure correct class
only N closes returned for <ticker> no_data_in_window (Tiingo answered thin)
fetch/timeout against Tiingo vendor_unreachable
HARInvalid: non-stationary fit / non-positive variance internal
Tiingo key not found internal (config)

A Tiingo outage therefore files as our own bug. The stated purpose of the change is to separate supply problems from ours; for one of the two vendors it does the opposite. FRED is classified correctly — mirror that treatment for HAR by having har_forecast.py emit a class, rather than inferring one from a prose string in the caller.

N1 — note: idempotency interaction

The additive fields mean a new-code record never == an old-code record, so the idempotency check will not suppress a re-run for a date already journaled by current code; a second record appears for that date. Defensible under "a changed observation is a new fact", but it should be stated in the concept rather than discovered in the journal.

What holds

  • approval: pending with verdict: proposed — correct gate, correctly open.
  • The proposal is genuinely off every deploy path (deploy.sh contains no reference to ops/). Verified.
  • missing strings preserved verbatim and in the original order. Verified by reading every branch, not by the shadow output.
  • Explicit not-adopted list, invalidation clause tied to base 735222a, provenance with a fetch sha and stale_after. This is the strongest governance work in the repo.

The design instinct is right. Fix B1, redo the proof per B2, split the classes per B3, and resubmit against the current base.