Skip to content

Verdict

REVISE — close. Everything the first review asked for was done, and the proof was re-run independently on hermes from commit 64081e8: 10/10 PASS, same forecast_rv30 0.1453, same staleness message, matching the committed proof-output.txt. The recorded output is honest.

Two defects remain. Neither is a B1-class catastrophe — the system fails closed correctly in both — but one files the blame wrong, which is the whole point of the change.

Confirmed fixed

  • B1 — open(JOURNAL, "a") restored. Proven the right way: case F reintroduces the defect and asserts the harness catches it (exit 1, journal 15->15, io.UnsupportedOperation: not writable). A proof able to demonstrate its own failure mode is exactly what was missing.
  • B2 — append enabled, sandbox seeded from the real regime.jsonl, comparison on the APPENDED records, exact commands printed. Credit for a detail that is easy to miss: the asof dates 2026-09-19 and 2026-09-26 are both Saturdays and therefore guaranteed absent from the journal, so the idempotency branch cannot suppress the write and the defective line is genuinely reached. That is thinking about how your own test could quietly pass.
  • B3 — solved structurally rather than patched: har_forecast.py emits {"error_class", "message"} itself; the caller never infers a class from prose. _fail preserves the stderr prose and SystemExit, so missing stays byte-compatible.
  • N1 — went past the brief. _norm_for_idempotency strips provenance.fetched_at before comparing; without it every re-run would append, since the timestamp always differs. The review flagged a symptom; the revision found the mechanism.

Independently verified: the only keys added are provenance and missing_errors.

R1-A — the change is NOT purely additive

forecast_model_version changes value:

deployed  531416d07752cc2b
proposed  142698820b90d053

strip_additives() removes it with a stated justification, which is honest, and the new value is correct — the file's bytes really did change. But this field is what establishes "these records came from the same model." A silent boundary at deploy reads, to anyone auditing the journal later, as a HAR model change when only error emission changed. The concept and the pending record both still assert that existing fields never change meaning, and that is now false.

Required: do not strip and forget. Declare it — in the concept and in the apply record — with both hashes, the deploy date, and the evidence already in hand that the computation is unchanged (forecast_rv30 and forecast_window_end identical on 2026-09-19). A reader hitting that boundary must find the explanation in the journal, not have to reconstruct it.

R1-B — an expired Tiingo key would be blamed on Tiingo

except Exception as e:   # network/timeout/HTTP against Tiingo -> vendor side
    _fail("vendor_unreachable", ...)

This catches every HTTPError, 401 and 403 included. A revoked, expired or mistyped Tiingo credential is ours, not the vendor's, and it is a far likelier real failure than Tiingo going down. The record would send the reader to a status page when the fix is rotating a key — the exact misdirection the taxonomy exists to prevent.

The proof demonstrates the bug rather than catching it. Case H2 injects a dummy key, receives <HTTPError 403: 'Forbidden'>, classifies it vendor_unreachable, and PASSES under the name "H2 har vendor failure -> vendor_unreachable JSON". The name asserts a semantic the injected fault does not support: the fault was an auth failure. This is the same species as B2, one level down — a test that passes without demonstrating what it claims. Worth sitting with, because it survived a revision written specifically to eliminate that pattern.

Required: classify by HTTP status, not by "any exception from the fetch call":

condition class
401, 403 internal (our credential)
connection refused, timeout, 5xx vendor_unreachable
404, 400 no_data_in_window (or internal if the request was malformed)

and change H2 to inject a connection failure (unroutable host or a closed port), so the test exercises the class it is named for.

R1-C — the third class is untested

no_data_in_window on the HAR path (thin close history, len(closes) < 100) is never exercised. The taxonomy claims three classes for HAR; the proof covers two. Add a case — a short --window, a truncated fixture, or a ticker with little history.

R1-D — trivial, but a trap

Case F computes ok and never uses it, and the expression has a precedence bug:

ok = r.returncode == 1 and after == before and "UnsupportedOperation" in r.stderr + r.stdout or (r.returncode == 1 and after == before)

and binds tighter than or, so this collapses to A and B — the stderr check is dead. The verdict line uses a correct expression, so nothing is wrong today. Delete it before someone reuses it believing it checks the message.

On resubmission

Fix R1-B, declare R1-A, add R1-C, delete R1-D, re-run the proof, and rebase onto the current base. R1-B is a small change; the test that goes with it is the part that matters.