RESOLVED 2026-10-01 — all nine items merged and deployed¶
Merged as PR #1 (
39c4c28, by the owner, 2026-10-01 11:46) fromfix/iv30-atm-review(eedb3c4+3444271). Live on both surfaces, md5c4fcf8f0; 64 tests pass on hermes. Apply record:2026-10-01:iv30-atm-review-fixes-merged.This file is closed. Take no action on its nine items — they are implemented. It is kept as the review of record, not as a work order.
Two things in it are superseded by what the fix pass found: - Item 9's framing was incomplete. It is listed as a test defect, which it was, but the pass also had to restore a REAL-quote check: asserting the arbiter's row from synthesised quotes verifies the CSV's arithmetic, not real quotes. Both tests now exist (
TestArbiterRegression,TestRealQuoteAgreement). - A tenth item existed and this file does not list it. The brief'sCORRECTION 2026-09-30specified step 5's replacement as cross-strike forward agreement; the first fix pass implemented only the forward band and leftiv30_forward_disagreementas a string nothing raised. hermes blocked the PR for it (REQUEST_CHANGES) and it is fixed in3444271. A check that cannot fire had been replaced by another check that cannot fire — see the PR discussion, which is the better record of that episode than this file.Downstream and still open: shadow runs (no agreement data yet) and the owner cutover gate (
a63ZvWo4GPXA8KtPw). Themapping_versionbump remains the owner's, per80-system-design/autonomy-authorization-2026-10-01.md.
REVIEW of tools/iv30_atm.py — accept the core, fix nine things before shadow¶
Reviewed at commit b8684c6. 21 tests, all passing, run independently.
Accepted, and the hard part is right¶
On the real 2026-09-29 quotes (Oct-16 8.38/8.40 and 8.74/8.76; Oct-30 13.35/13.39 and 10.91/10.98) the module returns 13.627 vol points against the arbiter's 13.689 for that date — 0.062 vol points apart. Parity forward 767.43 vs the arbiter's 767.44.
Also correct, and checked rather than assumed: atm_iv30 is genuinely pure (verified
by mutation test), the gate is untouched (git diff on regime_snapshot.py and
regime_config.json is empty), mapping_version appears nowhere, code_version is
computed at runtime with no literal hash in the source, total variance is interpolated
rather than vol, and the snapshot schema the CLI reads (as_of, spot.SPY.last,
metrics.SPY.expiration_ivs keys) matches the real files.
The pushback on step 5 is correct and is now my correction, not your defect. See
the CORRECTION section appended to 2026-09-30-BRIEF-hermes-derive-iv30.md: the
parity guard cannot fire by construction, I verified it independently (2,000 random
skews, 0 fires, max spread 5.8e-13), and the brief now specifies a
cross-strike forward agreement check instead, with a threshold measured off
5,013 real groups. Implement that replacement as part of this fix pass.
Must fix before this runs in shadow¶
1. _fail() mis-tags most failures as no_data_in_window¶
Every failure path gets error_class: "no_data_in_window". The taxonomy in
regime_snapshot.py has four classes — vendor_unreachable, no_data_in_window,
staleness, internal — and the distinction is the entire point: downstream needs
to tell "the vendor is down" from "our own maths refused the input".
Correct assignment:
| reason | class |
|---|---|
iv30_no_quotes, iv30_no_bracket, iv30_no_atm_pair |
no_data_in_window (correct today) |
iv30_quote_invalid, iv30_quote_crossed |
internal — data arrived and was malformed |
iv30_forward_out_of_range, iv30_forward_disagreement |
internal |
iv30_bisection_no_bracket, iv30_nonpositive_variance |
internal |
| MCP call raised / returned nothing | vendor_unreachable |
_fail() should take the class as an argument rather than hardcode one.
Also: _fail() sets "r": None even when r is known. Losing provenance on the
failure path is backwards — a failed round is exactly when you want to know what
rate was in play.
2. _assert_fail is too loose to catch #1¶
self.assertIn(info["missing"]["error_class"],
("no_data_in_window", "vendor_unreachable", "internal"))
This accepts any class, so it can never detect a mis-tag. Assert the specific expected class per test.
3. test_no_bisection_bracket asserts nothing¶
self.assertTrue(iv is None and info["missing"] is not None or iv is not None)
That is (A and B) or (not A) — true in three of four states, and the only failing
state (None, None) is unreachable because _fail() always populates missing. The
comments in the test body show it being written by trial and error.
The bisection-no-bracket path is therefore untested. Construct it deliberately:
pick F, K, T and a price above _black_undisc(F, K, T, 4.0, True), assert
_implied returns None, and separately assert atm_iv30 surfaces
iv30_bisection_no_bracket. Test the helper directly — that is much easier than
driving it through the whole function.
4. The receipt can attest quotes that were never used¶
In main(), receipt_syms.append(...) runs inside the per-leg loop, before the
if len(legs) == 2 check. A strike that yields only a call still contributes that
call's bid/ask to inputs.symbols, so the receipt names inputs that did not produce
the output. Build receipt_syms only from the strike that is finally accepted.
A receipt that misstates its inputs is worse than no receipt: it will be believed.
5. One malformed quote discards the whole day¶
The ladder steps outward when a leg is missing (continue), but a present
zero-bid or crossed quote returns _fail() immediately — so a bad quote at
round(spot) throws the day away even when a clean pair sits one dollar out.
It fails closed, so it is not dangerous, only wasteful: it will produce avoidable
no-IV30 days. Per-strike quote problems should continue to the next strike; fail
only when the ladder is exhausted, and then report which strikes were rejected and
why. Same treatment for iv30_forward_out_of_range and
iv30_bisection_no_bracket — they are properties of a candidate strike, not of the
day.
Keep iv30_no_bracket as an immediate failure. That one really is a property of the
day.
6. MAX_STRIKE_ATTEMPTS = 3 actually tries 5¶
ladder = [round(spot)]
for d in range(1, MAX_STRIKE_ATTEMPTS): # -> [S, S-1, S+1, S-2, S+2]
ladder += [round(spot) - d, round(spot) + d]
for k in ladder[: 2 * MAX_STRIKE_ATTEMPTS - 1]: # [:5] against a 5-element list
The slice is a no-op and the constant misleads by a factor of nearly two. Five
strikes is a fine choice — I have no objection to the behaviour — but name it for
what it does (STRIKE_LADDER_WIDTH = 2, or build the ladder to length
MAX_STRIKE_ATTEMPTS) and drop the dead slice.
7. rate_fallback is inferred from the value, not the source¶
"rate_fallback": (r == 0.0)
A genuine DGS1MO print of 0.00 — which happened repeatedly through 2020-21 — and an
explicit --r 0 both get labelled a fallback. Provenance must come from where the
number came from: pass the source into atm_iv30 (or pass a
rate_fallback boolean) and record r_source in info, not only in the receipt.
This is the same class of error as #4: info asserting something the inputs did not
establish, which the brief explicitly warned against and test_info_has_no_unjustified_values
does not currently test for.
8. --write rewrites the round snapshot in place, unguarded¶
with open(args.snapshot, "w") as f:
json.dump(snap, f, indent=1)
No backup, not atomic, and it re-serialises the entire capture at a formatting of its own choosing. A crash or a full disk mid-write destroys the round's snapshot, which is primary evidence.
Use the procedure this project already applies to jobs.json: write to a temp file in
the same directory, os.replace() it into place, and keep a .bak of the prior
content. Better still, write iv30_atm to a sidecar file and leave the capture
immutable — nothing else needs to mutate a snapshot after capture.
9. The regression test targets the wrong number — my fault¶
It hardcodes 13.7 and calls it "the arbiter's 2026-09-29 row". 13.7 is the
31-day leg I quoted in the brief; the arbiter's IV30 for that date is
13.689. The test passes either way, so nothing is broken, but it is not testing
what was asked.
This is on me: the brief pointed at backtest/out/iv30-spy-chain-inverted-2022-2026.csv,
which is off the deploy path and not present on hermes. Here is the row, inline,
so the test can be self-contained:
date=2026-09-29 iv30=0.136890 dte_lo=24 dte_hi=31 iv_lo=0.131793 iv_hi=0.137517
Assert within 0.5 vol points of 0.136890. Note the arbiter brackets with the
24-day expiry (Oct-23), not the 17-day one — with the full expiry calendar your
below[-1] will pick Oct-23 too, so supply the whole calendar in the test rather
than two expiries.
Minor¶
test_iv30_atm.py:59—e30 = ASOF.replace(year=2026)is dead, overwritten two lines later.test_shape_and_code_version:207—assertNotEqual(rec["code_version"], getattr(V, "SELF_HASH", None))compares againstNoneand is trivially true.- Two
ResourceWarnings from unclosed files inTestReceipt; usewith. - The near-expiry test quotes are commented "real capture 9/29 K=765" but are widened from the captured values (8.38/8.40 became 8.34/8.44). Either use the real numbers or drop the claim — this project has been bitten enough by comments that overstate their provenance.
One hazard from my side, worth knowing about¶
While calibrating the replacement check I found that tnx_history in VolSurfAE's
duckdb carries 720 NaN close_price rows — bond-market holidays such as Columbus
Day and Veterans Day, when equities trade but Treasury data is not published. My
arbiter's rate lookup returned NaN on those dates, DF became NaN, and 7 dates
silently vanished from the series (2022-10-10, 2022-11-11, 2023-10-09, 2024-10-14,
2024-11-11, 2025-10-13, 2025-11-11). It failed closed rather than fabricating a
number, but by accident, not design — and a NaN in a list makes sorted() return
garbage without raising, which is how I found it.
Fixed in backtest/iv30_from_chain.py; the arbiter now covers 1,189 dates and the
validation against the SVI series is unchanged (mean -0.001, median -0.025, sd 0.349
over 1,188 shared days).
Your fred_dgs1mo already handles FRED's "." placeholder, so you are not exposed to
the same bug. But assert math.isfinite() on r and on every computed F and
iv before they reach info. A non-finite number that propagates quietly is worse
than a failure, and it will not announce itself.