Skip to content

Summary

Amendment target: clause 4 of the EV contract and the edge_vs_rv output of all four attested computations.

Change: retire edge_vs_rv as a gate. Replace it with a layered gate — an absolute net-EV floor, a VRP-size edge test, and a forecast-uncertainty guard — and defer the conditional-vs-unconditional test until the calibration harness exists.

Why (Evidence)

The present test is algebraically vacuous

edge_vs_rv = ev_net(σ = garch_forecast) − ev_net(σ = rv_window) compares the same structure under two volatility assumptions. A short condor's EV is monotonically decreasing in σ, so the expression reduces to one bit:

edge_vs_rv > 0 ⟺ garch_forecast < rv_window

It carries no information about whether the trade is profitable. Reproduced on the blessed code (45-DTE SPX condor, 7150/7275 · 8075/8200):

IV RV forecast ev_net edge_vs_rv present gate
0.30 0.12 0.15 +26.96 −11.82 FAIL
0.12 0.20 0.15 −16.15 +16.52 PASS

Row 1 is an 18-vol-point variance risk premium with +27 EV — declined. Row 2 sells volatility below both realized and forecast, with negative EV — approved. The gate is inverted with respect to the premium-selling thesis: it rejects trades whenever the forecast sits above trailing realized, which is the ordinary state emerging from a calm stretch.

There is no profitability floor at all

No gate anywhere tests ev_net > 0. That absence, not the comparator's direction, is what admitted row 2. Any redefinition must close this first.

The day-1 precedent is misread

REVIEW.md §3 reads the 2026-09-09 decline as "the edge test rejects trades naive premium-selling would take." The receipt shows the mechanism was 0.11 > 0.0952. The decline was correct by coincidence, not by test. This amendment does not disturb that day's verdict — the event window deferred entry independently — but the precedent should not be cited as evidence that the edge test works.1

The Amended Rule

Layer 1 — profitability floor (absolute)

ev_net > 0

under the best available physical forecast, net of the full cost model. No candidate with ev_net ≤ 0 proceeds to any further gate, in any regime, under any operator exception. This is a never-rule in the sense of the risk function: it is not tradable through.

Layer 1 alone corrects every case in the table above.

Layer 2 — edge test (relative)

baseline_vrp = iv − forecast_rv  ≥  θ

with θ supplied by Amendment 2026-09-09/01. This answers the question clause 4 was written to ask — is volatility rich now — rather than the question the implementation asked.

Governance consequence: this makes θ load-bearing. Under v0 the EV gate "carries the full burden" and θ is advisory; under this amendment the θ calibration becomes a prerequisite for the edge test to function. The two amendments must therefore activate together, or Layer 2 activates with θ = 0 (sign only) as an explicit interim.

Layer 3 — forecast-uncertainty guard

iv − (forecast_rv + k · se_forecast)  >  0

where se_forecast is the standard error of the HAR fit and k is set by the operator (proposal: k = 1). A candidate whose edge does not survive the forecast's own error bar is not an edge; it is noise, and the honest response is reduced size or no trade.

This layer is expected to bind frequently. With the current daily squared-return RV proxy the HAR standard error is plausibly 2–4 volatility points, against an observed vrp_size series of 2.68 / 2.78 / 2.84 / 2.92 / 3.29 / 3.35. If Layer 3 rejects most days, that is the finding, not a defect — and it makes improving the RV proxy (range-based estimator from OHLC) a higher-priority piece of work than it currently is.

Deferred — Layer 4, conditional vs unconditional

ev_net(today, this regime)  −  mean ev_net(same archetype, all days)

This is the only layer that tests the framework's actual thesis: that regime gating adds value. It is the direct answer to option alphas being indistinguishable from zero since ~2012.2 It requires the calibration backtest harness — the same harness already blocking the θ calibration. One piece of work unblocks both; that is the argument for building it next.

Rejected Alternative — the dual-measure test

EV_PHYS − EV_RN > 0 was considered and rejected as a gate. It is attractive because ev_ladder.py already computes both terms, which would unify the two EV engines per the pilot reconciliation.

It fails because EV_RN ≈ −costs rather than ≈ 0 — the risk-neutral EV of a fairly priced structure is zero before costs — so subtracting it restores the cost drag and yields a gross-of-cost measure. Tested across the thin-edge regime, it passes trades that lose money:

IV forecast ev_net EV_PHYS − EV_RN
0.20 0.19 −5.09 +2.52 passes a loser
0.20 0.18 −1.97 +5.64 passes a loser
0.25 0.235 −6.44 +3.12 passes a loser
0.18 0.17 −3.91 +2.74 passes a loser
0.16 0.15 −2.66 +2.92 passes a loser

Five of five. EV_RN retains its correct role, which is the one the pilot already gives it: a validation signature that marks and model agree (EV_RN ≈ 0 at entry for a straddle), not an edge test.

Activation Conditions

  • [x] Layer 1 (ev_net > 0) — independent of the others; implemented 2026-09-17, emitted as gate_ev_net_positive by all four computations
  • [ ] Layer 2 — requires θ from Amendment 2026-09-09/01, or an explicit interim θ = 0
  • [ ] Layer 3 — requires se_forecast exposed by tools/har_forecast.py
  • [ ] Layer 4 — requires the calibration backtest harness
  • [ ] Human approval recorded in the decision log

Until activation of Layers 2–4, edge_vs_rv remains emitted but is no longer a gate; it is retained as a diagnostic so the historical receipt series stays comparable. Layer 1 is in force from its implementation date.

Documented Prediction

Per the learning function, this amendment predicts:

  1. Layer 1 changes no historical verdict — the day-1 condor was declined on the event window regardless. Falsifiable against the receipt series.
  2. Layer 2 admits trades the present gate rejects, specifically rich-VRP days where the forecast sits above trailing realized. Expect more eligible days, not fewer.
  3. Layer 3 removes most of them again. If it does not, se_forecast is understated and the HAR proxy needs replacing before θ can mean anything.

Prediction 2 and prediction 3 pull in opposite directions by design: the pair is the falsifiable content.

Version Ledger

Version Edge rule Status
v0 (current) edge_vs_rv > 0 superseded as a gate; retained as diagnostic
v1 (this proposal) Layer 1 in force; Layers 2–4 staged partial — Layer 1 active 2026-09-17

References

  • Review findings 2026-09-17, blocker B1: ../references/receipts/2026-09-17-review-findings.md
  • Chicago Fed Working Paper 2025-17: https://www.chicagofed.org/-/media/publications/working-papers/2025/wp2025-17.pdf
  • Corsi (2009), "A Simple Approximate Long-Memory Model of Realized Volatility," Journal of Financial Econometrics 7(2) — the HAR-RV forecast whose error Layer 3 guards against.

  1. Reproduced by executing the blessed computations; see the findings document. ↩

  2. Chicago Fed WP 2025-17. ↩