BRIEF for hermes — derive SPY ATM IV30 from quotes¶
Owner decision 2026-09-30: iv30 means SPY ATM IV30, and hermes derives it
itself — no daily dependency on ai-rig.
Read 95-inbox/2026-09-30-iv30-source-selection.md before starting. Its conclusion
governs this build: do not use metrics.SPY.iv30, do not use expiration_ivs, and
do not use per-leg option_quotes[...].volatility. All three were measured against
a chain-inversion arbiter. The first two are VIX-style variance-swap constructions
and sit +2.67 and +2.70 vol points above ATM. The third violates put-call parity
by a median 2.71 and up to 14.53 vol points between a call and a put on the same
strike and expiry. Reaching for any of them relabels the defect instead of fixing it.
Build on bid/ask. Quotes are the only values in this feed that are not somebody's model output.
Scope¶
Produce IV30 alongside the existing proxy, into the round snapshot. Do not wire
it into the regime gate. The cutover flips the VRP sign on 40.7% of days and
relabels history; it needs a mapping_version bump and its own amendment, and that
is the owner's call, not part of this task. Shadow first — that also gives us live
agreement data between your number and the arbiter's.
Deliverable¶
options-system-wiki/tools/iv30_atm.py, plus tools/test_iv30_atm.py.
Structure it the way regime_snapshot.py was refactored: a pure function doing the
mathematics, with no I/O, no clock, no argv, and no mutation of its inputs, and a
thin CLI around it. That is what makes it testable, and this is a fail-closed input to
a trading gate.
def atm_iv30(spot, expiries, quotes, asof, r=0.0):
"""-> (iv30 | None, info)
spot : float, SPY last/mid
expiries : [date] -- the available expiry calendar
quotes : {(expiry, strike, 'C'|'P'): (bid, ask)}
asof : date -- passed in, never read from the clock
r : float -- continuous annual rate
info carries, always: the two expiries chosen, their DTEs, the ATM strike,
the parity forward, per-expiry ATM IV, the call/put IV spread at each
expiry, r and whether it was a fallback, and on failure a structured
`missing` entry in the existing missing_errors taxonomy.
"""
Algorithm¶
- Bracket 30 calendar days.
T = (expiry - asof).days / 365. Need one expiry withT <= 30/365and one withT >= 30/365. Never extrapolate past the nearest available expiry to reach 30 days — that is the rule the SVI work already follows, and 1 date in 1,188 legitimately failed it. - ATM strike: the strike nearest
spotthat has both a call and a put quoted at that expiry. Use one strike for both legs — parity only holds at a shared strike. - Forward by put-call parity:
F = K + (C_mid - P_mid) / DF,DF = exp(-rT). Do not use spot as the forward; that is the likely cause of the vendor field's parity violation. - Invert Black on the undiscounted mid (
mid / DF), by bisection on[1e-4, 4.0], call and put separately. - Cross-check, do not average blindly: if
|iv_call - iv_put| > 1.0vol point at an expiry, that expiry is unusable — fail closed withiv30_parity_inconsistent. Average only when they agree. This is the check that would have caught the vendor field, so it belongs in our own code too. - Interpolate total variance linearly in T:
w = iv^2 * T, interpolate toT30 = 30/365, returnsqrt(w / T30). Total variance, not implied vol — linear interpolation of total variance is what keeps the term structure arbitrage-free.
Fail-closed conditions¶
Return None with a structured missing entry — never a substituted or
best-effort number — when: an expiry is missing on either side of 30 days; the ATM
strike has only one leg; bid <= 0, ask <= 0, or ask < bid; the bisection does
not bracket; the parity forward is outside 0.5*spot .. 2.0*spot; or the call/put IV
spread exceeds 1.0 vol point.
The rate¶
Prefer FRED DGS1MO — correct tenor, and you already have the FRED fetch pattern
in regime_snapshot.py. If it is unavailable, use r = 0.0 and record
rate_fallback: true in info; do not fail the computation for it. That is a
deliberate exception to fail-closed and it is justified by measurement, not
convenience: re-running the full 1,182-day arbiter at r = 0 moves the series by
-0.05 vol points (sd 0.02, worst case -0.52). Record it so the exception stays
visible.
Data acquisition (the CLI side)¶
- spot: already captured.
- expiry calendar: the keys of
metrics.SPY.expiration_ivs. Use the keys only — the values are the unusable index. Free, and nothing new to fetch. - quotes:
tastytrade_get_quotewithinstrument_type: "Equity Option", symbols built by the existingocc()helper incapture_snapshot.py. 4 symbols per round — ATM call and put at each of the two bracketing expiries. If the first strike you try is missing a leg, step to the next nearest strike; cap at 3 attempts per expiry, then fail closed.
Provenance¶
Emit an attested receipt in the established shape — {computation, inputs, outputs,
code_version} with code_version = sha256(file). inputs must name the 4 symbols,
their bid/ask, spot, r and its source. Do not embed the file's own hash as a
literal in the file — that is self-referential and was removed from
har_forecast.py for that reason; compute it at runtime.
Tests (TDD — write them first, watch them fail)¶
Cover at least: exact-30-day expiry needs no interpolation; asymmetric bracket
(24d/31d) interpolates in total variance and not in vol — assert the two differ
and that you match the variance answer; missing near expiry → fail closed; missing far
expiry → fail closed; ATM strike with only a call → steps to the next strike; crossed
quote (ask < bid) → fail closed; call/put IV spread of 3 points → iv30_parity_inconsistent;
absurd parity forward → fail closed; r fallback sets the flag and still returns a
number; info never contains a value the inputs did not justify; and the pure function
does not mutate its quotes argument.
One regression test with real numbers, because we have the arbiter's answer:
2026-09-29, SPY, spot 764.20, Oct-16 and Oct-30 expiries, ATM K=765, Oct-30 call mid
13.37 / put mid 10.945 → parity forward ≈ 767.44 → 31-day ATM IV ≈ 13.7%. The
arbiter's interpolated IV30 for 2026-09-29 is in
backtest/out/iv30-spy-chain-inverted-2022-2026.csv. Assert you land within 0.5 vol
points of that row.
Acceptance¶
tools/test_iv30_atm.pypasses, and you watched each test fail first.- The round snapshot carries
iv30_atmwith its receipt, next to the existing proxy. - A journal record with the code_version, the shadow-mode framing, and an explicit statement that the gate is unchanged.
- Do not touch
regime_config.json,regime_snapshot.py's gate logic, ormapping_version. If you think the cutover should happen, say so in the record and leave it to the owner.
One thing to push back on if I have it wrong¶
The 1.0 vol point parity tolerance in step 5 is my judgement, not a measurement. If real ATM SPY quotes routinely disagree by more than that for a benign reason — wide spreads near the close, a stale leg — say so with numbers from the rounds you observe, and propose a tolerance you can defend. Do not widen it silently to make the check pass.
CORRECTION 2026-09-30 — step 5 was vacuous, and hermes proved it¶
This section supersedes step 5 above. The original text is left in place unedited, because hermes built against it and the record should show what was actually asked.
What I got wrong¶
Step 5 told you to reject an expiry when |iv_call - iv_put| > 1.0 vol point, and
called it "the check that would have caught the vendor field". That check cannot
fire. Step 3 derives the forward from put-call parity at the ATM strike, which
forces C_mid - P_mid = DF*(F - K) exactly. Two legs priced consistently with a
forward, inverted at that same forward, return the identical implied vol — that is
what parity means. A fat or stale leg does not widen the spread; it moves F.
hermes identified this, kept the guard defensively, and documented it in
test_parity_guard_cannot_fire_by_construction with the evidence. I confirmed it
independently before accepting the pushback: 2,000 random call/put price
combinations, 0 fires, maximum observed spread 5.8e-13 vol points.
The guard is not harmful — it costs two subtractions and would become meaningful if
anyone later computes F independently per leg. But as written it is decoration,
and a decorative check is worse than none because it reads as protection. That is
the same standard applied to the adapter's trusted-uid refusal; it applies to my own
specifications too.
Note also what the original claim got wrong about the vendor: Tastytrade's per-leg IVs disagreed because they were not built from a parity-consistent forward. Our construction cannot reproduce that failure, so no check on our own two legs could ever have "caught" it. The check I asked for was answering a question we do not have.
Step 5, replacement: cross-strike forward agreement¶
The forward is one number per expiry. Every strike's parity equation must imply
the same F. Dispersion across strikes therefore is a real quote-integrity signal
— a stale, wide or mis-parsed leg shows up immediately — and unlike the original it
is not satisfied by construction.
You already fetch a strike ladder, so this costs nothing extra beyond keeping the quotes you discarded.
- At the chosen expiry, compute the parity forward
F_iat each of the strikes in the ladder for which both legs are quoted (you need at least 3; if fewer are available, recordfwd_agreement: insufficient_strikesand proceed — do not fail). - Convert the dispersion to the unit that matters, its impact on ATM vol:
impact_vol_pts = (max(F_i) - min(F_i)) / (spot * sqrt(T)) * 100 - Fail closed with
iv30_forward_disagreementwhenimpact_vol_pts > 1.0. - Record
fwd_spread_usd,fwd_impact_vol_ptsandn_strikes_usedininfowhether or not it fires. A check whose normal-range value is invisible cannot be audited.
Where the 1.0 comes from — measured, not chosen¶
Over the real SPY chain, all 5,013 (date x expiry) groups at 20-40 DTE from 2022-01-03 to 2026-09-29, using the 5 strikes nearest spot:
| cross-strike forward spread | p50 | p90 | p99 | p99.9 | max |
|---|---|---|---|---|---|
| dollars | 0.056 | 0.130 | 0.583 | 1.041 | 1.334 |
| vol points of ATM impact | 0.037 | 0.091 | 0.304 | 0.659 | 0.773 |
A 1.0 vol-point threshold never fires in four and a half years of real quotes. That is the property a fail-closed guard needs: no false positives on good data, so a fire is genuinely anomalous. It is not a number I liked the look of — the last one was, and it was wrong.
If you find it firing in live rounds, report the distribution you observe and propose a threshold from it. Do not widen it to make a round pass.
Also tighten: the forward band in step 3¶
FWD_LO, FWD_HI = 0.5, 2.0 admits a forward 50% below or 100% above spot. Nothing
that wrong could arise from a real quote; it would only ever catch a catastrophe.
Measured F/spot - 1 over the same 5,013 groups:
| p0.1 | p1 | p50 | p99 | p99.9 | min | max |
|---|---|---|---|---|---|---|
| -0.725% | -0.512% | +0.211% | +0.708% | +1.114% | -0.802% | +1.359% |
Use +/-3%. That is more than double the worst observation in the history, so it will not fire spuriously, while still catching a mis-parsed strike, a wrong expiry or a units error — the blunders that actually happen.