Skip to content

Choosing the IV30 source: every ready-made field is the wrong quantity

Owner decision 2026-09-30: iv30 means SPY ATM IV30, and hermes derives it itself (no daily coupling to ai-rig).

Acting on that produced an unwelcome finding. hermes already receives three candidate IV fields every round, for free, from the Tastytrade MCP. All three are unusable, and one of them is internally incoherent. This note records the arbitration, because the cheapest path here is also the wrong one and that needs to be on the record before someone takes it.

The arbiter

backtest/iv30_from_chain.py computes ATM IV30 with no fitted surface and no vendor index, from the raw SPY chain in VolSurfAE's duckdb (option_chains, 11.4M SPY rows, 2022-01-03 to 2026-09-29):

  1. per (date, expiry), take the strike nearest spot with both legs quoted
  2. forward by put-call parity, F = K + (C - P)/DF
  3. Black inversion by bisection on the undiscounted price, call and put separately, then averaged -- the two must agree under parity, so their spread is a built-in check
  4. bracket 30 calendar days and interpolate total variance linearly in T

1,182 dates scored, 0 unbracketed. Median 0.1499.

r comes from tnx_history (10Y), which is the wrong tenor. Rather than assert that this does not matter, the script reruns at r = 0: the whole series moves -0.05 vol points (sd 0.02, worst case -0.52). It does not matter.

Result 1 -- the canonical series is confirmed, by an independent method

The SVI-fit series (2026-09-30:canonical-iv30-measurement, read at k=0 off VolSurfAE's fitted parameters) versus this chain inversion, over 1,181 shared days:

vol points
mean -0.00
median -0.03
sd 0.35
90th pct of abs diff 0.51
max abs diff 2.94

Two derivations sharing no code and no intermediate artifact agree to a hundredth of a vol point on average. The measurement stands, and so do the two corrections filed against the theta evidence and the long-vol finding.

This check was worth running for a specific reason: the fits in feature_store/svi_params are near-degenerate (b ~ 1.0, rho ~ +0.93, sigma ~ 0.01), and a positive rho is the wrong sign for an equity index. At k=0 such a fit evaluates as a difference of two nearly equal terms, which is exactly where cancellation error would hide. It evidently does not bite here -- but the SVI store should not be trusted at k=0 on the strength of r_squared alone, and there are several superseded copies of that store (svi_params_refit, svi_params_refit_25al, two .pre_*_backup) that a future reader could easily pick up by mistake.

Result 2 -- all three ready-made fields are wrong, by the same ~2.6 points

Against the arbiter, on the 12 overlapping round-snapshot dates:

candidate mean median sd range
VIX proxy (what the gate uses today) +2.62 +2.51 1.06 -0.14 .. +9.92
metrics.SPY.iv30 (Tastytrade) +2.67 +2.60 0.38 +2.17 .. +3.38
expiration_ivs interpolated to 30d +2.70 +2.61 0.35 +2.30 .. +3.38

Tastytrade's implied-volatility-index is a VIX-style variance-swap construction, not an ATM read, and its per-expiration implied-volatility is the same kind of object. That is why they agree with VIX to within 0.1 vol points and miss ATM by 2.6.

Their agreement with VIX is not corroboration. It briefly looked like two independent sources contradicting the canonical series; they are two instances of the same methodology. Swapping the VIX proxy for either Tastytrade field would change nothing about the defect -- it would relabel it.

Result 3 -- per-leg volatility violates put-call parity

option_quotes[sym].volatility looked like the ideal input: a genuine per-contract IV at a strike we choose. It is not usable. Over 52 call/put pairs sharing a strike and an expiry, where parity forces near-equal IV:

| |call IV - put IV|, vol points | |---| | mean 3.67, median 2.71, sd 3.73, min 0.15, max 14.53 |

Example, 2026-09-29 SPY Oct-16 K=765: call 14.11%, put 11.73%. Averaging the pair happens to land near the right answer, which is the signature of a forward/carry misspecification -- but a field with an unexplained 2.7-point median internal contradiction is not something to build a fail-closed gate on.

What hermes should do instead

Derive it, from quotes rather than from anyone's vol field. Everything needed is already in place on hermes:

need already available
SPY spot capture_snapshot captures it
the expiry calendar keys of metrics.SPY.expiration_ivs -- free, and reliable
ATM option quotes tastytrade_get_quote on Equity Option, symbols built by the existing occ() helper
the maths math.erf -- stdlib, no dependency

Per round: pick the two expiries bracketing 30 days, build 4 OCC symbols (ATM call + put at each), fetch the quotes, parity-forward, invert, interpolate total variance. Four extra quote symbols per round. No ThetaData, no credential to relocate, no ai-rig coupling, no new third-party package on the trading host.

Bid/ask are quotes, not derived values -- they are the one thing in this data that is not somebody's model output. That is the reason to build on them.

Not decided here

Switching the gate's iv30 from the VIX proxy to SPY ATM IV30 flips the VRP sign on 40.7% of days and relabels history. That is a change in gate behaviour, not an added parameter, so the long_vol_archetype precedent (deliberately no mapping_version bump, following max_input_age_business_days) does not cover it. This one should carry a mapping_version bump and its own amendment record. The owner's decision authorises the definition; it does not by itself authorise a silent cutover. Flagged for the owner, not assumed.