Skip to content

BRIEF for hermes — derive SPY ATM IV30 from quotes

Owner decision 2026-09-30: iv30 means SPY ATM IV30, and hermes derives it itself — no daily dependency on ai-rig.

Read 95-inbox/2026-09-30-iv30-source-selection.md before starting. Its conclusion governs this build: do not use metrics.SPY.iv30, do not use expiration_ivs, and do not use per-leg option_quotes[...].volatility. All three were measured against a chain-inversion arbiter. The first two are VIX-style variance-swap constructions and sit +2.67 and +2.70 vol points above ATM. The third violates put-call parity by a median 2.71 and up to 14.53 vol points between a call and a put on the same strike and expiry. Reaching for any of them relabels the defect instead of fixing it.

Build on bid/ask. Quotes are the only values in this feed that are not somebody's model output.

Scope

Produce IV30 alongside the existing proxy, into the round snapshot. Do not wire it into the regime gate. The cutover flips the VRP sign on 40.7% of days and relabels history; it needs a mapping_version bump and its own amendment, and that is the owner's call, not part of this task. Shadow first — that also gives us live agreement data between your number and the arbiter's.

Deliverable

options-system-wiki/tools/iv30_atm.py, plus tools/test_iv30_atm.py.

Structure it the way regime_snapshot.py was refactored: a pure function doing the mathematics, with no I/O, no clock, no argv, and no mutation of its inputs, and a thin CLI around it. That is what makes it testable, and this is a fail-closed input to a trading gate.

def atm_iv30(spot, expiries, quotes, asof, r=0.0):
    """-> (iv30 | None, info)

    spot     : float, SPY last/mid
    expiries : [date]           -- the available expiry calendar
    quotes   : {(expiry, strike, 'C'|'P'): (bid, ask)}
    asof     : date             -- passed in, never read from the clock
    r        : float            -- continuous annual rate

    info carries, always: the two expiries chosen, their DTEs, the ATM strike,
    the parity forward, per-expiry ATM IV, the call/put IV spread at each
    expiry, r and whether it was a fallback, and on failure a structured
    `missing` entry in the existing missing_errors taxonomy.
    """

Algorithm

  1. Bracket 30 calendar days. T = (expiry - asof).days / 365. Need one expiry with T <= 30/365 and one with T >= 30/365. Never extrapolate past the nearest available expiry to reach 30 days — that is the rule the SVI work already follows, and 1 date in 1,188 legitimately failed it.
  2. ATM strike: the strike nearest spot that has both a call and a put quoted at that expiry. Use one strike for both legs — parity only holds at a shared strike.
  3. Forward by put-call parity: F = K + (C_mid - P_mid) / DF, DF = exp(-rT). Do not use spot as the forward; that is the likely cause of the vendor field's parity violation.
  4. Invert Black on the undiscounted mid (mid / DF), by bisection on [1e-4, 4.0], call and put separately.
  5. Cross-check, do not average blindly: if |iv_call - iv_put| > 1.0 vol point at an expiry, that expiry is unusable — fail closed with iv30_parity_inconsistent. Average only when they agree. This is the check that would have caught the vendor field, so it belongs in our own code too.
  6. Interpolate total variance linearly in T: w = iv^2 * T, interpolate to T30 = 30/365, return sqrt(w / T30). Total variance, not implied vol — linear interpolation of total variance is what keeps the term structure arbitrage-free.

Fail-closed conditions

Return None with a structured missing entry — never a substituted or best-effort number — when: an expiry is missing on either side of 30 days; the ATM strike has only one leg; bid <= 0, ask <= 0, or ask < bid; the bisection does not bracket; the parity forward is outside 0.5*spot .. 2.0*spot; or the call/put IV spread exceeds 1.0 vol point.

The rate

Prefer FRED DGS1MO — correct tenor, and you already have the FRED fetch pattern in regime_snapshot.py. If it is unavailable, use r = 0.0 and record rate_fallback: true in info; do not fail the computation for it. That is a deliberate exception to fail-closed and it is justified by measurement, not convenience: re-running the full 1,182-day arbiter at r = 0 moves the series by -0.05 vol points (sd 0.02, worst case -0.52). Record it so the exception stays visible.

Data acquisition (the CLI side)

  • spot: already captured.
  • expiry calendar: the keys of metrics.SPY.expiration_ivs. Use the keys only — the values are the unusable index. Free, and nothing new to fetch.
  • quotes: tastytrade_get_quote with instrument_type: "Equity Option", symbols built by the existing occ() helper in capture_snapshot.py. 4 symbols per round — ATM call and put at each of the two bracketing expiries. If the first strike you try is missing a leg, step to the next nearest strike; cap at 3 attempts per expiry, then fail closed.

Provenance

Emit an attested receipt in the established shape — {computation, inputs, outputs, code_version} with code_version = sha256(file). inputs must name the 4 symbols, their bid/ask, spot, r and its source. Do not embed the file's own hash as a literal in the file — that is self-referential and was removed from har_forecast.py for that reason; compute it at runtime.

Tests (TDD — write them first, watch them fail)

Cover at least: exact-30-day expiry needs no interpolation; asymmetric bracket (24d/31d) interpolates in total variance and not in vol — assert the two differ and that you match the variance answer; missing near expiry → fail closed; missing far expiry → fail closed; ATM strike with only a call → steps to the next strike; crossed quote (ask < bid) → fail closed; call/put IV spread of 3 points → iv30_parity_inconsistent; absurd parity forward → fail closed; r fallback sets the flag and still returns a number; info never contains a value the inputs did not justify; and the pure function does not mutate its quotes argument.

One regression test with real numbers, because we have the arbiter's answer: 2026-09-29, SPY, spot 764.20, Oct-16 and Oct-30 expiries, ATM K=765, Oct-30 call mid 13.37 / put mid 10.945 → parity forward ≈ 767.44 → 31-day ATM IV ≈ 13.7%. The arbiter's interpolated IV30 for 2026-09-29 is in backtest/out/iv30-spy-chain-inverted-2022-2026.csv. Assert you land within 0.5 vol points of that row.

Acceptance

  • tools/test_iv30_atm.py passes, and you watched each test fail first.
  • The round snapshot carries iv30_atm with its receipt, next to the existing proxy.
  • A journal record with the code_version, the shadow-mode framing, and an explicit statement that the gate is unchanged.
  • Do not touch regime_config.json, regime_snapshot.py's gate logic, or mapping_version. If you think the cutover should happen, say so in the record and leave it to the owner.

One thing to push back on if I have it wrong

The 1.0 vol point parity tolerance in step 5 is my judgement, not a measurement. If real ATM SPY quotes routinely disagree by more than that for a benign reason — wide spreads near the close, a stale leg — say so with numbers from the rounds you observe, and propose a tolerance you can defend. Do not widen it silently to make the check pass.


CORRECTION 2026-09-30 — step 5 was vacuous, and hermes proved it

This section supersedes step 5 above. The original text is left in place unedited, because hermes built against it and the record should show what was actually asked.

What I got wrong

Step 5 told you to reject an expiry when |iv_call - iv_put| > 1.0 vol point, and called it "the check that would have caught the vendor field". That check cannot fire. Step 3 derives the forward from put-call parity at the ATM strike, which forces C_mid - P_mid = DF*(F - K) exactly. Two legs priced consistently with a forward, inverted at that same forward, return the identical implied vol — that is what parity means. A fat or stale leg does not widen the spread; it moves F.

hermes identified this, kept the guard defensively, and documented it in test_parity_guard_cannot_fire_by_construction with the evidence. I confirmed it independently before accepting the pushback: 2,000 random call/put price combinations, 0 fires, maximum observed spread 5.8e-13 vol points.

The guard is not harmful — it costs two subtractions and would become meaningful if anyone later computes F independently per leg. But as written it is decoration, and a decorative check is worse than none because it reads as protection. That is the same standard applied to the adapter's trusted-uid refusal; it applies to my own specifications too.

Note also what the original claim got wrong about the vendor: Tastytrade's per-leg IVs disagreed because they were not built from a parity-consistent forward. Our construction cannot reproduce that failure, so no check on our own two legs could ever have "caught" it. The check I asked for was answering a question we do not have.

Step 5, replacement: cross-strike forward agreement

The forward is one number per expiry. Every strike's parity equation must imply the same F. Dispersion across strikes therefore is a real quote-integrity signal — a stale, wide or mis-parsed leg shows up immediately — and unlike the original it is not satisfied by construction.

You already fetch a strike ladder, so this costs nothing extra beyond keeping the quotes you discarded.

  1. At the chosen expiry, compute the parity forward F_i at each of the strikes in the ladder for which both legs are quoted (you need at least 3; if fewer are available, record fwd_agreement: insufficient_strikes and proceed — do not fail).
  2. Convert the dispersion to the unit that matters, its impact on ATM vol: impact_vol_pts = (max(F_i) - min(F_i)) / (spot * sqrt(T)) * 100
  3. Fail closed with iv30_forward_disagreement when impact_vol_pts > 1.0.
  4. Record fwd_spread_usd, fwd_impact_vol_pts and n_strikes_used in info whether or not it fires. A check whose normal-range value is invisible cannot be audited.

Where the 1.0 comes from — measured, not chosen

Over the real SPY chain, all 5,013 (date x expiry) groups at 20-40 DTE from 2022-01-03 to 2026-09-29, using the 5 strikes nearest spot:

cross-strike forward spread p50 p90 p99 p99.9 max
dollars 0.056 0.130 0.583 1.041 1.334
vol points of ATM impact 0.037 0.091 0.304 0.659 0.773

A 1.0 vol-point threshold never fires in four and a half years of real quotes. That is the property a fail-closed guard needs: no false positives on good data, so a fire is genuinely anomalous. It is not a number I liked the look of — the last one was, and it was wrong.

If you find it firing in live rounds, report the distribution you observe and propose a threshold from it. Do not widen it to make a round pass.

Also tighten: the forward band in step 3

FWD_LO, FWD_HI = 0.5, 2.0 admits a forward 50% below or 100% above spot. Nothing that wrong could arise from a real quote; it would only ever catch a catastrophe. Measured F/spot - 1 over the same 5,013 groups:

p0.1 p1 p50 p99 p99.9 min max
-0.725% -0.512% +0.211% +0.708% +1.114% -0.802% +1.359%

Use +/-3%. That is more than double the worst observation in the history, so it will not fire spuriously, while still catching a mis-parsed strike, a wrong expiry or a units error — the blunders that actually happen.