Summary¶
Amendment target: clause 4 of the EV contract
and the edge_vs_rv output of all four
attested computations.
Change: retire edge_vs_rv as a gate. Replace it with a layered gate —
an absolute net-EV floor, a VRP-size edge test, and a forecast-uncertainty
guard — and defer the conditional-vs-unconditional test until the calibration
harness exists.
Why (Evidence)¶
The present test is algebraically vacuous¶
edge_vs_rv = ev_net(σ = garch_forecast) − ev_net(σ = rv_window) compares the
same structure under two volatility assumptions. A short condor's EV is
monotonically decreasing in σ, so the expression reduces to one bit:
edge_vs_rv > 0⟺garch_forecast < rv_window
It carries no information about whether the trade is profitable. Reproduced on the blessed code (45-DTE SPX condor, 7150/7275 · 8075/8200):
| IV | RV | forecast | ev_net |
edge_vs_rv |
present gate |
|---|---|---|---|---|---|
| 0.30 | 0.12 | 0.15 | +26.96 | −11.82 | FAIL |
| 0.12 | 0.20 | 0.15 | −16.15 | +16.52 | PASS |
Row 1 is an 18-vol-point variance risk premium with +27 EV — declined. Row 2 sells volatility below both realized and forecast, with negative EV — approved. The gate is inverted with respect to the premium-selling thesis: it rejects trades whenever the forecast sits above trailing realized, which is the ordinary state emerging from a calm stretch.
There is no profitability floor at all¶
No gate anywhere tests ev_net > 0. That absence, not the comparator's
direction, is what admitted row 2. Any redefinition must close this first.
The day-1 precedent is misread¶
REVIEW.md §3 reads the 2026-09-09 decline as "the edge test
rejects trades naive premium-selling would take." The receipt shows the
mechanism was 0.11 > 0.0952. The decline was correct by coincidence, not by
test. This amendment does not disturb that day's verdict — the event window
deferred entry independently — but the precedent should not be cited as
evidence that the edge test works.1
The Amended Rule¶
Layer 1 — profitability floor (absolute)¶
ev_net > 0
under the best available physical forecast, net of the full cost model. No
candidate with ev_net ≤ 0 proceeds to any further gate, in any regime, under
any operator exception. This is a never-rule in the sense of the
risk function: it is not tradable through.
Layer 1 alone corrects every case in the table above.
Layer 2 — edge test (relative)¶
baseline_vrp = iv − forecast_rv ≥ θ
with θ supplied by Amendment 2026-09-09/01. This answers the question clause 4 was written to ask — is volatility rich now — rather than the question the implementation asked.
Governance consequence: this makes θ load-bearing. Under v0 the EV gate "carries the full burden" and θ is advisory; under this amendment the θ calibration becomes a prerequisite for the edge test to function. The two amendments must therefore activate together, or Layer 2 activates with θ = 0 (sign only) as an explicit interim.
Layer 3 — forecast-uncertainty guard¶
iv − (forecast_rv + k · se_forecast) > 0
where se_forecast is the standard error of the HAR fit and k is set by the
operator (proposal: k = 1). A candidate whose edge does not survive the
forecast's own error bar is not an edge; it is noise, and the honest response
is reduced size or no trade.
This layer is expected to bind frequently. With the current daily
squared-return RV proxy the HAR standard error is plausibly 2–4 volatility
points, against an observed vrp_size series of 2.68 / 2.78 / 2.84 / 2.92 /
3.29 / 3.35. If Layer 3 rejects most days, that is the finding, not a defect
— and it makes improving the RV proxy (range-based estimator from OHLC) a
higher-priority piece of work than it currently is.
Deferred — Layer 4, conditional vs unconditional¶
ev_net(today, this regime) − mean ev_net(same archetype, all days)
This is the only layer that tests the framework's actual thesis: that regime gating adds value. It is the direct answer to option alphas being indistinguishable from zero since ~2012.2 It requires the calibration backtest harness — the same harness already blocking the θ calibration. One piece of work unblocks both; that is the argument for building it next.
Rejected Alternative — the dual-measure test¶
EV_PHYS − EV_RN > 0 was considered and rejected as a gate. It is attractive
because ev_ladder.py already computes both terms, which would unify the two EV
engines per the
pilot reconciliation.
It fails because EV_RN ≈ −costs rather than ≈ 0 — the risk-neutral EV of a
fairly priced structure is zero before costs — so subtracting it restores the
cost drag and yields a gross-of-cost measure. Tested across the thin-edge
regime, it passes trades that lose money:
| IV | forecast | ev_net |
EV_PHYS − EV_RN |
|
|---|---|---|---|---|
| 0.20 | 0.19 | −5.09 | +2.52 | passes a loser |
| 0.20 | 0.18 | −1.97 | +5.64 | passes a loser |
| 0.25 | 0.235 | −6.44 | +3.12 | passes a loser |
| 0.18 | 0.17 | −3.91 | +2.74 | passes a loser |
| 0.16 | 0.15 | −2.66 | +2.92 | passes a loser |
Five of five. EV_RN retains its correct role, which is the one the pilot
already gives it: a validation signature that marks and model agree
(EV_RN ≈ 0 at entry for a straddle), not an edge test.
Activation Conditions¶
- [x] Layer 1 (
ev_net > 0) — independent of the others; implemented 2026-09-17, emitted asgate_ev_net_positiveby all four computations - [ ] Layer 2 — requires θ from Amendment 2026-09-09/01, or an explicit interim θ = 0
- [ ] Layer 3 — requires
se_forecastexposed bytools/har_forecast.py - [ ] Layer 4 — requires the calibration backtest harness
- [ ] Human approval recorded in the decision log
Until activation of Layers 2–4, edge_vs_rv remains emitted but is no longer
a gate; it is retained as a diagnostic so the historical receipt series stays
comparable. Layer 1 is in force from its implementation date.
Documented Prediction¶
Per the learning function, this amendment predicts:
- Layer 1 changes no historical verdict — the day-1 condor was declined on the event window regardless. Falsifiable against the receipt series.
- Layer 2 admits trades the present gate rejects, specifically rich-VRP days where the forecast sits above trailing realized. Expect more eligible days, not fewer.
- Layer 3 removes most of them again. If it does not,
se_forecastis understated and the HAR proxy needs replacing before θ can mean anything.
Prediction 2 and prediction 3 pull in opposite directions by design: the pair is the falsifiable content.
Version Ledger¶
| Version | Edge rule | Status |
|---|---|---|
| v0 (current) | edge_vs_rv > 0 |
superseded as a gate; retained as diagnostic |
| v1 (this proposal) | Layer 1 in force; Layers 2–4 staged | partial — Layer 1 active 2026-09-17 |
References¶
- Review findings 2026-09-17, blocker B1: ../references/receipts/2026-09-17-review-findings.md
- Chicago Fed Working Paper 2025-17: https://www.chicagofed.org/-/media/publications/working-papers/2025/wp2025-17.pdf
- Corsi (2009), "A Simple Approximate Long-Memory Model of Realized Volatility," Journal of Financial Econometrics 7(2) — the HAR-RV forecast whose error Layer 3 guards against.