Skip to content

RESOLVED 2026-10-01 — all nine items merged and deployed

Merged as PR #1 (39c4c28, by the owner, 2026-10-01 11:46) from fix/iv30-atm-review (eedb3c4 + 3444271). Live on both surfaces, md5 c4fcf8f0; 64 tests pass on hermes. Apply record: 2026-10-01:iv30-atm-review-fixes-merged.

This file is closed. Take no action on its nine items — they are implemented. It is kept as the review of record, not as a work order.

Two things in it are superseded by what the fix pass found: - Item 9's framing was incomplete. It is listed as a test defect, which it was, but the pass also had to restore a REAL-quote check: asserting the arbiter's row from synthesised quotes verifies the CSV's arithmetic, not real quotes. Both tests now exist (TestArbiterRegression, TestRealQuoteAgreement). - A tenth item existed and this file does not list it. The brief's CORRECTION 2026-09-30 specified step 5's replacement as cross-strike forward agreement; the first fix pass implemented only the forward band and left iv30_forward_disagreement as a string nothing raised. hermes blocked the PR for it (REQUEST_CHANGES) and it is fixed in 3444271. A check that cannot fire had been replaced by another check that cannot fire — see the PR discussion, which is the better record of that episode than this file.

Downstream and still open: shadow runs (no agreement data yet) and the owner cutover gate (a63ZvWo4GPXA8KtPw). The mapping_version bump remains the owner's, per 80-system-design/autonomy-authorization-2026-10-01.md.

REVIEW of tools/iv30_atm.py — accept the core, fix nine things before shadow

Reviewed at commit b8684c6. 21 tests, all passing, run independently.

Accepted, and the hard part is right

On the real 2026-09-29 quotes (Oct-16 8.38/8.40 and 8.74/8.76; Oct-30 13.35/13.39 and 10.91/10.98) the module returns 13.627 vol points against the arbiter's 13.689 for that date — 0.062 vol points apart. Parity forward 767.43 vs the arbiter's 767.44.

Also correct, and checked rather than assumed: atm_iv30 is genuinely pure (verified by mutation test), the gate is untouched (git diff on regime_snapshot.py and regime_config.json is empty), mapping_version appears nowhere, code_version is computed at runtime with no literal hash in the source, total variance is interpolated rather than vol, and the snapshot schema the CLI reads (as_of, spot.SPY.last, metrics.SPY.expiration_ivs keys) matches the real files.

The pushback on step 5 is correct and is now my correction, not your defect. See the CORRECTION section appended to 2026-09-30-BRIEF-hermes-derive-iv30.md: the parity guard cannot fire by construction, I verified it independently (2,000 random skews, 0 fires, max spread 5.8e-13), and the brief now specifies a cross-strike forward agreement check instead, with a threshold measured off 5,013 real groups. Implement that replacement as part of this fix pass.


Must fix before this runs in shadow

1. _fail() mis-tags most failures as no_data_in_window

Every failure path gets error_class: "no_data_in_window". The taxonomy in regime_snapshot.py has four classes — vendor_unreachable, no_data_in_window, staleness, internal — and the distinction is the entire point: downstream needs to tell "the vendor is down" from "our own maths refused the input".

Correct assignment:

reason class
iv30_no_quotes, iv30_no_bracket, iv30_no_atm_pair no_data_in_window (correct today)
iv30_quote_invalid, iv30_quote_crossed internal — data arrived and was malformed
iv30_forward_out_of_range, iv30_forward_disagreement internal
iv30_bisection_no_bracket, iv30_nonpositive_variance internal
MCP call raised / returned nothing vendor_unreachable

_fail() should take the class as an argument rather than hardcode one.

Also: _fail() sets "r": None even when r is known. Losing provenance on the failure path is backwards — a failed round is exactly when you want to know what rate was in play.

2. _assert_fail is too loose to catch #1

self.assertIn(info["missing"]["error_class"],
              ("no_data_in_window", "vendor_unreachable", "internal"))

This accepts any class, so it can never detect a mis-tag. Assert the specific expected class per test.

3. test_no_bisection_bracket asserts nothing

self.assertTrue(iv is None and info["missing"] is not None or iv is not None)

That is (A and B) or (not A) — true in three of four states, and the only failing state (None, None) is unreachable because _fail() always populates missing. The comments in the test body show it being written by trial and error.

The bisection-no-bracket path is therefore untested. Construct it deliberately: pick F, K, T and a price above _black_undisc(F, K, T, 4.0, True), assert _implied returns None, and separately assert atm_iv30 surfaces iv30_bisection_no_bracket. Test the helper directly — that is much easier than driving it through the whole function.

4. The receipt can attest quotes that were never used

In main(), receipt_syms.append(...) runs inside the per-leg loop, before the if len(legs) == 2 check. A strike that yields only a call still contributes that call's bid/ask to inputs.symbols, so the receipt names inputs that did not produce the output. Build receipt_syms only from the strike that is finally accepted.

A receipt that misstates its inputs is worse than no receipt: it will be believed.

5. One malformed quote discards the whole day

The ladder steps outward when a leg is missing (continue), but a present zero-bid or crossed quote returns _fail() immediately — so a bad quote at round(spot) throws the day away even when a clean pair sits one dollar out.

It fails closed, so it is not dangerous, only wasteful: it will produce avoidable no-IV30 days. Per-strike quote problems should continue to the next strike; fail only when the ladder is exhausted, and then report which strikes were rejected and why. Same treatment for iv30_forward_out_of_range and iv30_bisection_no_bracket — they are properties of a candidate strike, not of the day.

Keep iv30_no_bracket as an immediate failure. That one really is a property of the day.

6. MAX_STRIKE_ATTEMPTS = 3 actually tries 5

ladder = [round(spot)]
for d in range(1, MAX_STRIKE_ATTEMPTS):   # -> [S, S-1, S+1, S-2, S+2]
    ladder += [round(spot) - d, round(spot) + d]
for k in ladder[: 2 * MAX_STRIKE_ATTEMPTS - 1]:   # [:5] against a 5-element list

The slice is a no-op and the constant misleads by a factor of nearly two. Five strikes is a fine choice — I have no objection to the behaviour — but name it for what it does (STRIKE_LADDER_WIDTH = 2, or build the ladder to length MAX_STRIKE_ATTEMPTS) and drop the dead slice.

7. rate_fallback is inferred from the value, not the source

"rate_fallback": (r == 0.0)

A genuine DGS1MO print of 0.00 — which happened repeatedly through 2020-21 — and an explicit --r 0 both get labelled a fallback. Provenance must come from where the number came from: pass the source into atm_iv30 (or pass a rate_fallback boolean) and record r_source in info, not only in the receipt.

This is the same class of error as #4: info asserting something the inputs did not establish, which the brief explicitly warned against and test_info_has_no_unjustified_values does not currently test for.

8. --write rewrites the round snapshot in place, unguarded

with open(args.snapshot, "w") as f:
    json.dump(snap, f, indent=1)

No backup, not atomic, and it re-serialises the entire capture at a formatting of its own choosing. A crash or a full disk mid-write destroys the round's snapshot, which is primary evidence.

Use the procedure this project already applies to jobs.json: write to a temp file in the same directory, os.replace() it into place, and keep a .bak of the prior content. Better still, write iv30_atm to a sidecar file and leave the capture immutable — nothing else needs to mutate a snapshot after capture.

9. The regression test targets the wrong number — my fault

It hardcodes 13.7 and calls it "the arbiter's 2026-09-29 row". 13.7 is the 31-day leg I quoted in the brief; the arbiter's IV30 for that date is 13.689. The test passes either way, so nothing is broken, but it is not testing what was asked.

This is on me: the brief pointed at backtest/out/iv30-spy-chain-inverted-2022-2026.csv, which is off the deploy path and not present on hermes. Here is the row, inline, so the test can be self-contained:

date=2026-09-29  iv30=0.136890  dte_lo=24  dte_hi=31  iv_lo=0.131793  iv_hi=0.137517

Assert within 0.5 vol points of 0.136890. Note the arbiter brackets with the 24-day expiry (Oct-23), not the 17-day one — with the full expiry calendar your below[-1] will pick Oct-23 too, so supply the whole calendar in the test rather than two expiries.


Minor

  • test_iv30_atm.py:59 — e30 = ASOF.replace(year=2026) is dead, overwritten two lines later.
  • test_shape_and_code_version:207 — assertNotEqual(rec["code_version"], getattr(V, "SELF_HASH", None)) compares against None and is trivially true.
  • Two ResourceWarnings from unclosed files in TestReceipt; use with.
  • The near-expiry test quotes are commented "real capture 9/29 K=765" but are widened from the captured values (8.38/8.40 became 8.34/8.44). Either use the real numbers or drop the claim — this project has been bitten enough by comments that overstate their provenance.

One hazard from my side, worth knowing about

While calibrating the replacement check I found that tnx_history in VolSurfAE's duckdb carries 720 NaN close_price rows — bond-market holidays such as Columbus Day and Veterans Day, when equities trade but Treasury data is not published. My arbiter's rate lookup returned NaN on those dates, DF became NaN, and 7 dates silently vanished from the series (2022-10-10, 2022-11-11, 2023-10-09, 2024-10-14, 2024-11-11, 2025-10-13, 2025-11-11). It failed closed rather than fabricating a number, but by accident, not design — and a NaN in a list makes sorted() return garbage without raising, which is how I found it.

Fixed in backtest/iv30_from_chain.py; the arbiter now covers 1,189 dates and the validation against the SVI series is unchanged (mean -0.001, median -0.025, sd 0.349 over 1,188 shared days).

Your fred_dgs1mo already handles FRED's "." placeholder, so you are not exposed to the same bug. But assert math.isfinite() on r and on every computed F and iv before they reach info. A non-finite number that propagates quietly is worse than a failure, and it will not announce itself.