Choosing the IV30 source: every ready-made field is the wrong quantity¶
Owner decision 2026-09-30: iv30 means SPY ATM IV30, and hermes derives it
itself (no daily coupling to ai-rig).
Acting on that produced an unwelcome finding. hermes already receives three candidate IV fields every round, for free, from the Tastytrade MCP. All three are unusable, and one of them is internally incoherent. This note records the arbitration, because the cheapest path here is also the wrong one and that needs to be on the record before someone takes it.
The arbiter¶
backtest/iv30_from_chain.py computes ATM IV30 with no fitted surface and no vendor
index, from the raw SPY chain in VolSurfAE's duckdb (option_chains, 11.4M SPY rows,
2022-01-03 to 2026-09-29):
- per (date, expiry), take the strike nearest spot with both legs quoted
- forward by put-call parity,
F = K + (C - P)/DF - Black inversion by bisection on the undiscounted price, call and put separately, then averaged -- the two must agree under parity, so their spread is a built-in check
- bracket 30 calendar days and interpolate total variance linearly in T
1,182 dates scored, 0 unbracketed. Median 0.1499.
r comes from tnx_history (10Y), which is the wrong tenor. Rather than assert that
this does not matter, the script reruns at r = 0: the whole series moves
-0.05 vol points (sd 0.02, worst case -0.52). It does not matter.
Result 1 -- the canonical series is confirmed, by an independent method¶
The SVI-fit series (2026-09-30:canonical-iv30-measurement, read at k=0 off
VolSurfAE's fitted parameters) versus this chain inversion, over 1,181 shared days:
| vol points | |
|---|---|
| mean | -0.00 |
| median | -0.03 |
| sd | 0.35 |
| 90th pct of abs diff | 0.51 |
| max abs diff | 2.94 |
Two derivations sharing no code and no intermediate artifact agree to a hundredth of a vol point on average. The measurement stands, and so do the two corrections filed against the theta evidence and the long-vol finding.
This check was worth running for a specific reason: the fits in
feature_store/svi_params are near-degenerate (b ~ 1.0, rho ~ +0.93, sigma ~
0.01), and a positive rho is the wrong sign for an equity index. At k=0 such a fit
evaluates as a difference of two nearly equal terms, which is exactly where
cancellation error would hide. It evidently does not bite here -- but the SVI store
should not be trusted at k=0 on the strength of r_squared alone, and there are
several superseded copies of that store (svi_params_refit, svi_params_refit_25al,
two .pre_*_backup) that a future reader could easily pick up by mistake.
Result 2 -- all three ready-made fields are wrong, by the same ~2.6 points¶
Against the arbiter, on the 12 overlapping round-snapshot dates:
| candidate | mean | median | sd | range |
|---|---|---|---|---|
| VIX proxy (what the gate uses today) | +2.62 | +2.51 | 1.06 | -0.14 .. +9.92 |
metrics.SPY.iv30 (Tastytrade) |
+2.67 | +2.60 | 0.38 | +2.17 .. +3.38 |
expiration_ivs interpolated to 30d |
+2.70 | +2.61 | 0.35 | +2.30 .. +3.38 |
Tastytrade's implied-volatility-index is a VIX-style variance-swap construction,
not an ATM read, and its per-expiration implied-volatility is the same kind of
object. That is why they agree with VIX to within 0.1 vol points and miss ATM by 2.6.
Their agreement with VIX is not corroboration. It briefly looked like two independent sources contradicting the canonical series; they are two instances of the same methodology. Swapping the VIX proxy for either Tastytrade field would change nothing about the defect -- it would relabel it.
Result 3 -- per-leg volatility violates put-call parity¶
option_quotes[sym].volatility looked like the ideal input: a genuine per-contract
IV at a strike we choose. It is not usable. Over 52 call/put pairs sharing a strike
and an expiry, where parity forces near-equal IV:
| |call IV - put IV|, vol points | |---| | mean 3.67, median 2.71, sd 3.73, min 0.15, max 14.53 |
Example, 2026-09-29 SPY Oct-16 K=765: call 14.11%, put 11.73%. Averaging the pair happens to land near the right answer, which is the signature of a forward/carry misspecification -- but a field with an unexplained 2.7-point median internal contradiction is not something to build a fail-closed gate on.
What hermes should do instead¶
Derive it, from quotes rather than from anyone's vol field. Everything needed is already in place on hermes:
| need | already available |
|---|---|
| SPY spot | capture_snapshot captures it |
| the expiry calendar | keys of metrics.SPY.expiration_ivs -- free, and reliable |
| ATM option quotes | tastytrade_get_quote on Equity Option, symbols built by the existing occ() helper |
| the maths | math.erf -- stdlib, no dependency |
Per round: pick the two expiries bracketing 30 days, build 4 OCC symbols (ATM call + put at each), fetch the quotes, parity-forward, invert, interpolate total variance. Four extra quote symbols per round. No ThetaData, no credential to relocate, no ai-rig coupling, no new third-party package on the trading host.
Bid/ask are quotes, not derived values -- they are the one thing in this data that is not somebody's model output. That is the reason to build on them.
Not decided here¶
Switching the gate's iv30 from the VIX proxy to SPY ATM IV30 flips the VRP sign
on 40.7% of days and relabels history. That is a change in gate behaviour, not an
added parameter, so the long_vol_archetype precedent (deliberately no
mapping_version bump, following max_input_age_business_days) does not cover
it. This one should carry a mapping_version bump and its own amendment record.
The owner's decision authorises the definition; it does not by itself authorise a
silent cutover. Flagged for the owner, not assumed.