Verdict¶
REVISE — close. Everything the first review asked for was done, and the proof was
re-run independently on hermes from commit 64081e8: 10/10 PASS, same
forecast_rv30 0.1453, same staleness message, matching the committed
proof-output.txt. The recorded output is honest.
Two defects remain. Neither is a B1-class catastrophe — the system fails closed correctly in both — but one files the blame wrong, which is the whole point of the change.
Confirmed fixed¶
- B1 —
open(JOURNAL, "a")restored. Proven the right way: case F reintroduces the defect and asserts the harness catches it (exit 1, journal 15->15,io.UnsupportedOperation: not writable). A proof able to demonstrate its own failure mode is exactly what was missing. - B2 — append enabled, sandbox seeded from the real
regime.jsonl, comparison on the APPENDED records, exact commands printed. Credit for a detail that is easy to miss: the asof dates 2026-09-19 and 2026-09-26 are both Saturdays and therefore guaranteed absent from the journal, so the idempotency branch cannot suppress the write and the defective line is genuinely reached. That is thinking about how your own test could quietly pass. - B3 — solved structurally rather than patched:
har_forecast.pyemits{"error_class", "message"}itself; the caller never infers a class from prose._failpreserves the stderr prose andSystemExit, somissingstays byte-compatible. - N1 — went past the brief.
_norm_for_idempotencystripsprovenance.fetched_atbefore comparing; without it every re-run would append, since the timestamp always differs. The review flagged a symptom; the revision found the mechanism.
Independently verified: the only keys added are provenance and missing_errors.
R1-A — the change is NOT purely additive¶
forecast_model_version changes value:
deployed 531416d07752cc2b
proposed 142698820b90d053
strip_additives() removes it with a stated justification, which is honest, and the new
value is correct — the file's bytes really did change. But this field is what
establishes "these records came from the same model." A silent boundary at deploy reads,
to anyone auditing the journal later, as a HAR model change when only error emission
changed. The concept and the pending record both still assert that existing fields never
change meaning, and that is now false.
Required: do not strip and forget. Declare it — in the concept and in the apply
record — with both hashes, the deploy date, and the evidence already in hand that the
computation is unchanged (forecast_rv30 and forecast_window_end identical on
2026-09-19). A reader hitting that boundary must find the explanation in the journal,
not have to reconstruct it.
R1-B — an expired Tiingo key would be blamed on Tiingo¶
except Exception as e: # network/timeout/HTTP against Tiingo -> vendor side
_fail("vendor_unreachable", ...)
This catches every HTTPError, 401 and 403 included. A revoked, expired or mistyped
Tiingo credential is ours, not the vendor's, and it is a far likelier real failure
than Tiingo going down. The record would send the reader to a status page when the fix is
rotating a key — the exact misdirection the taxonomy exists to prevent.
The proof demonstrates the bug rather than catching it. Case H2 injects a dummy key,
receives <HTTPError 403: 'Forbidden'>, classifies it vendor_unreachable, and PASSES
under the name "H2 har vendor failure -> vendor_unreachable JSON". The name asserts a
semantic the injected fault does not support: the fault was an auth failure. This is the
same species as B2, one level down — a test that passes without demonstrating what it
claims. Worth sitting with, because it survived a revision written specifically to
eliminate that pattern.
Required: classify by HTTP status, not by "any exception from the fetch call":
| condition | class |
|---|---|
| 401, 403 | internal (our credential) |
| connection refused, timeout, 5xx | vendor_unreachable |
| 404, 400 | no_data_in_window (or internal if the request was malformed) |
and change H2 to inject a connection failure (unroutable host or a closed port), so the test exercises the class it is named for.
R1-C — the third class is untested¶
no_data_in_window on the HAR path (thin close history, len(closes) < 100) is never
exercised. The taxonomy claims three classes for HAR; the proof covers two. Add a case —
a short --window, a truncated fixture, or a ticker with little history.
R1-D — trivial, but a trap¶
Case F computes ok and never uses it, and the expression has a precedence bug:
ok = r.returncode == 1 and after == before and "UnsupportedOperation" in r.stderr + r.stdout or (r.returncode == 1 and after == before)
and binds tighter than or, so this collapses to A and B — the stderr check is
dead. The verdict line uses a correct expression, so nothing is wrong today. Delete it
before someone reuses it believing it checks the message.
On resubmission¶
Fix R1-B, declare R1-A, add R1-C, delete R1-D, re-run the proof, and rebase onto the current base. R1-B is a small change; the test that goes with it is the part that matters.