Review findings — 2026-09-17¶
Review of the framework as it runs in production. Findings are open unless marked otherwise; this repository's initial commit is the "before" baseline, so each fix lands as its own reviewable diff.
Every claim below was reproduced by running the code, not read off the source.
Blockers¶
B1 — The EV gate is vacuous, and inverted relative to the VRP thesis (Layer 1 FIXED 2026-09-17; Layers 2–4 proposed)¶
edge_vs_rv = ev_net(σ=garch) − ev_net(σ=rv) (references/computations/ev_iron_condor.py:206).
Both terms price the same structure under two vol assumptions. A short condor's
EV is monotonically decreasing in σ, so the gate reduces to one bit:
edge_vs_rv > 0⟺garch_forecast < rv_window
Reproduced:
| IV | RV | HAR | ev_net |
edge_vs_rv |
Gate |
|---|---|---|---|---|---|
| 0.30 | 0.12 | 0.15 | +26.96 | −11.82 | FAIL |
| 0.12 | 0.20 | 0.15 | −16.15 | +16.52 | PASS |
Row 1: an 18-vol-point VRP with +27 EV — declined. Row 2: selling vol below both realized and forecast, negative EV — approved. The gate passes money-losers and rejects the best VRP trades whenever the forecast sits above trailing realized, which is the normal state coming out of a calm stretch.
REVIEW.md §3 reads the day-1 decline as "the edge test rejects trades naive
premium-selling would take." What actually happened was 0.11 > 0.0952.
The root cause is narrower than the comparator's direction: no gate anywhere
tested ev_net > 0. That absence is what admitted row 2.
Status: resolved by Amendment 2026-09-17/01,
which retires edge_vs_rv as a gate (retaining it as a diagnostic) and replaces
it with a layered test:
| Layer | Test | Status |
|---|---|---|
| 1 | ev_net > 0, net of costs |
in force 2026-09-17 |
| 2 | iv − forecast_rv ≥ θ |
pending θ calibration |
| 3 | edge survives the forecast's standard error | pending se_forecast |
| 4 | conditional EV beats the unconditional archetype | pending the backtest harness |
Layer 1 is implemented: all four computations emit gate_ev_net_positive, and
the attester reports a separate ev_gate derived from the mandatory ev_net —
backward compatible with pre-amendment receipts — exiting non-zero when either
check fails. Verified 5/5 on the cases above; tampered-flag case negative-tested.
The dual-measure alternative EV_PHYS − EV_RN was tested and rejected:
EV_RN ≈ −costs, not ≈ 0, so subtracting it restores the cost drag and passes
losing trades (5 of 5 in the thin-edge regime).
B2 — The attester does not attest¶
references/attesters/ev-binding.py validates key names, types, path binding and
sha256(computation file). It never re-runs the computation. Taking the real
day-1 receipt and flipping ev_net 12.73 → 9999, edge_vs_rv −5.75 → +9998,
inputs untouched:
{"verdict": "pass", "reasons": []} exit=0
Charter principle 2 is "No attestation, no proposal"; the stated failure mode it prevents is "LLM arithmetic, non-reproducible numbers." A fabricated receipt passes.
Fix: import the module, call compute_ev(**receipt["inputs"]), compare with
receipt["outputs"] within tolerance. Roughly ten lines.
Related: code_version = sha256(file) means any edit invalidates every historical
receipt with no recovery path. Retain versioned source, not just its hash.
B3 — max_loss reports a risk-free position that can lose $17k¶
ev_iron_condor.py:194 takes min() of the two wings. With put wing 200 and call
wing 25:
reported max_loss: -0.68 <- negative: "no maximum loss"
true max_loss: 174.32 <- $17,432 per contract
max_loss feeds position sizing and the 20% heat cap. ev_butterfly.py rejects
asymmetric wings; the condor accepts and misreports them.
Fix: max(), and subtract costs.
B4 — Lookahead in the live regime journal¶
tools/regime_snapshot.py:64 — vix3m_from_cboe() takes no --asof and fetches
the current delayed quote. Visible in the journal:
date: 2026-09-11 vix_date: 2026-09-10 vix3m_asof: 2026-09-13 01:52:51
The 9/11 term ratio was computed from 9/13 data. That ratio was 0.9591 — 1% from
flipping flat→contango and changing the label. regime-definition.md calls
ex-ante computability non-negotiable. Any point-in-time replay used to calibrate θ
inherits this.
Also d["data"]["close"] or d["data"]["current_price"] mixes prior-close and
intraday depending on run time.
Majors¶
M1 — The framework's most conservative rule is not implemented.
transition_lookback_days: 5 appears in tools/regime_config.json and nowhere in
the code. The docs define a transition as "the confirmed label changed within the
last N days" — a cooling-off window. The code defines it as raw ≠ confirmed
now, so the day a flip confirms, transition goes False and new risk resumes
immediately. confirmation_days is equally decorative: line 147 compares it to
len(recent) but only tests recent[-1] == recent[-2], hardcoding 2.
M2 — Same-day re-runs can falsely confirm a regime label. A duplicate append
makes recent[-1] == recent[-2] trivially true. Live evidence: regime-state.json
lists 2026-09-16 three times. In this instance 9/15 and 9/16 were both
calm-contango-vrp+, so the 9/17 transition was legitimate — right answer,
unsound mechanism. The log records this as "minor idempotency"; it gates Gate 1
eligibility.
M3 — Fail-closed catches missing data, never stale data. vix_from_fred
scans back 45 days and returns the newest value found, reporting
data_missing: false. A FRED outage yields a three-week-old VIX indistinguishable
from a fresh one. Same for vix3m_ts and har["window_end"]. Needs a max-age
assertion per input.
M4 — outcomes.jsonl does not exist on either host. The learning function
exists to close forecast → reality, and the concordance quorum needs ≥3 concordant
outcome events. No file, no writer, Hermes loop stage 10 undone. The amendment
process is unreachable by its primary path.
M5 — The deployed EV engine is unattested. ev_ladder.py won the
consolidation but has no AC concept and emits ladder JSON the attester cannot
parse, while hermes-cron-prompt-v2.md is status: stable and applied. The live
daily round contradicts charter principle 2 continuously. Should be a recorded
exception with an end date, not implicit.
M6 — Journal fork (healed in this repo's initial commit). hermes held
regime.jsonl (9 records) and no decisions.jsonl; the workstation held
decisions.jsonl (17 records) and a 2-record regime.jsonl. Neither was a
superset. Merged here; single-writer rule now documented in README.md.
M7 — Six of seven framework documents are absent from the production host.
hermes has only decision-function.md (an older revision, missing the
regime_config.json canonical reference). risk-function.md — the never-rules,
the frozen v0 risk parameters, the standing tail rules — is not present where
the agent runs. log.md 2026-09-16 claims "canonical wiki has all 88-framework
docs on both machines"; that claim is false for hermes. Fixed by deploy.sh.
M8 — No actor field in the decision log. Records assert "owner approved" inside agent-written free text. For the one inviolable rule — the human gate — there is no machine-checkable distinction between a human approved and the agent wrote that a human approved. The OKF spec has an actor convention the decision schema does not use.
M9 — 82 of 285 relative provenance links are broken (one ../ short); four
files over-corrected to ../../../../. log.md records fixing these in
sources[] frontmatter; body footnotes were missed. Provenance is the wiki's
headline feature.
M10 — Documented prod config does not match deployed prod config.
hermes-cron-prompt-v2.md and the archived cron-prompt-v2.*-applied.txt copies
all reference /home/ysakakibara/...; the live jobs.json uses /home/hermes/...
and is 7003 chars against v2.2's archived 6744. A second job, report-builder, is
undocumented. Both live prompts are now exported to ops/cron/.
Moderate¶
- HAR fit quality. Squared daily close-to-close returns are a very noisy RV
proxy — errors-in-variables will attenuate
b_daily; OLS on variance levels lets a few crash days dominate. Corsi fits intraday RV, usually in logs. Tiingo already returns OHLC, so a Garman–Klass estimator is ~5× more efficient for free. Nothing checksb_d + b_w + b_m < 1before iterating 22 steps. - Two incompatible VRP definitions. Snapshot:
VIX − HAR_forecast. EV code:iv − rv_window(trailing). Both called VRP; θ is calibrated against the first. - θ is source-conditional. VIX is a variance-swap strike, not ATM IV30. When ThetaData IV30 lands (V1b) the VRP level shifts systematically by ~1–2 vol points and any θ calibrated now becomes invalid.
- Default
slippage_pct = 0.15— $527 on a $3,511 SPX condor credit — contradicts the charter's own "zero commissions, penny-wide SPX" row. - EV models hold-to-expiry while the pilot's headline lesson is "the exit path IS the edge" (SPY 765 straddle, EV_PHYS −21.9%). The attested number does not describe the trade being proposed.
- Drift drops the dividend yield (
μ = r + ERP, ~1.2%/yr high on a price index) ande_valueis an undiscounted terminal expectation added to a PV credit. Both small, both systematic. - Tiingo key passed in the URL query string (
tools/har_forecast.py) — use theAuthorization: Tokenheader. REVIEW.mdredefines the calibration gate. Open decision 2 says calibration "awaits sufficient VRP_size series," citing three live prints. The amendment's own plan requires a 2016→present ex-ante series with purged walk-forward and the 2018/2020/2024 stress set — runnable today, sincehar_forecast.py --asofalready loops. The real blocker is the absent backtest harness.- The decision log is ~90% self-administration. Of 17 records, 2 concern trades; the rest are cron prompts, env hardening, ledger corrections, migration verification.
What holds up¶
60-regimes/amendment-2026-09-09-vrp-size-threshold.mddisqualifies its own supporting evidence, downgrades its quorum claim, and pre-commits to "or the honest conclusion that no θ works." Rare discipline.- Fail-closed as a first-class concept; missing data as a state, not an assumption.
regime-detection-methods.md— filtered vs smoothed, published detection-delay distributions, ±20% threshold perturbation.- Forced resolution, versioned rules, no mid-session amendment: the right structural answer to LLM drift.
- The research brief was allowed to undercut the project's own premise (VRP alphas ≈ 0 post-2012) and drive a redesign rather than a defence.
Suggested order¶
- B2 — make the attester re-execute the computation
- B3 —
max()of wings - B4 + M3 — thread
--asofto VIX3M; max-age guards on all inputs - M1 + M2 — implement
transition_lookback_days, honourconfirmation_days, keyrecentby date so re-runs are idempotent - M4 — create
outcomes.jsonland its writer - M9 — fix the links (one
sedpass) - ~~B1 — redesign the edge test~~ — Layer 1 done; Layers 2–4 await θ
calibration,
se_forecast, and the backtest harness
One piece of work unblocks the most: the calibration backtest harness gates both the θ calibration (Amendment 2026-09-09/01) and Layer 4 of the edge test. It is the highest-leverage item on this list.
1–6 are roughly a week. None of the six framework concepts should be promoted to
stable before they land: the drafts currently describe a more rigorous system
than the one running.
Appendix — Empirical Basis for the Amendment 2026-09-09/01 Review¶
Added 2026-09-17. Reproduce with tools/vrp_series.py (steps 1–2 of the
amendment's own calibration plan), then the analysis below. Inputs are the same
proxies the production tool uses — VIX as IV30, SPY close-to-close RV — so these
figures describe what this rule does with this pipeline. Effect ranking is
robust; exact percentages inherit the proxies' limitations.
Series: 3,396 ex-ante daily observations, 2013-03-18 → 2026-09-16. HAR-RV refit on trailing 756 observations at each date, data ≤ t only.
θ = 3.0 is substantially a VIX-level filter¶
corr(VIX, VRP_size) = +0.476; slope +0.288 vol points of VRP per VIX point.
| VIX band | days | pass @ θ=3.0 | mean VRP |
|---|---|---|---|
| <13 | 675 | 0.0% | 0.09 |
| 13–16 | 993 | 14.0% | 1.34 |
| 16–20 | 851 | 39.0% | 1.93 |
| 20–25 | 484 | 52.7% | 3.38 |
| 25–30 | 235 | 77.4% | 5.69 |
| >30 | 149 | 80.5% | 7.39 |
Against the mapping it serves: 18.7% of calm days (VIX<20, premium selling eligible) versus 64.2% of elevated days (premium selling forbidden). θ is most permissive where the mapping forbids trading and most restrictive where it permits it. Over 675 days at VIX<13, not one day would have been eligible.
Undisclosed consequence: eligible calm days fall from 187/yr to 35/yr, an 81% reduction.
θ sits inside the forecast noise¶
Day-over-day VRP_size change: sd 2.88 vol points. The threshold is crossed
every 8.2 days; 22.4% of all days fall within ±1.0 of it. The six live
observations (2.68 / 2.78 / 2.84 / 2.92 / 3.29 / 3.35) span 0.67 — under a
quarter of one day's typical move. They sample noise around θ rather than
converging on it. No hysteresis is applied to this axis, unlike the VIX and term
axes, contrary to regime-definition.md.
The calibration cannot produce the number it asks for¶
A 45-DTE trade occupies ~31 trading days, so 2013–2026 yields ~109 independent trades: ~17 in bucket 0–1, ~32 in 1–3, ~21 in 3–6, ~11 in >6. The mandatory stress set is ~1 independent trade per episode, 3 total — a sanity check, not a statistical test. Selecting "the smallest bucket boundary at which net edge turns positive" across those buckets is threshold-mining; the plan requires a multiple-testing control but names none.
The forecast breaks exactly where the stress set lives¶
All 9 implausible forecasts in 13 years fall inside stress windows:
2020-03-10 124% 2020-03-13 559% 2020-03-24 165%
2020-03-11 116% 2020-03-16 1008% 2025-04-09 133%
2020-03-12 466% 2020-03-17 461% 2025-04-10 119%
The mandatory stress-set validation would run on inputs like a 1008% annualized vol forecast. Missing stationarity check on the iterated HAR path.
Scale-invariance helps but does not close it¶
VRP_ratio = VIX/forecast − 1 at a matched 30.4% pass rate drops
corr(VIX, ·) from +0.476 to +0.292 and flattens the bands
(1.2 / 24.2 / 38.0 / 44.4 / 60.9 / 67.8%) — but they remain monotone. VRP is
genuinely regime-dependent, so no single global θ is correct; θ must be
conditioned on the regime band. The calibration plan's strata are on VRP_size
buckets when the variable requiring stratification is the regime itself.
Correction — re-run under the HAR validity guard (2026-09-17, later)¶
The figures above were computed on a hand-filtered series (days with a
forecast above 100% or below 2% dropped). tools/har_forecast.py now carries a
principled validity guard that refuses 20 days — explosive fits (companion
spectral radius ≥ 1) and stable-but-sign-flipped fits that predict negative
variance, which the hand filter kept (e.g. 2015-08-31 at 8.3%). Re-run on the
guarded series (3,376 days, 20 excluded):
| Figure | hand-filtered (above) | guarded | |
|---|---|---|---|
| corr(VIX, VRP_size) | +0.476 | +0.498 | conclusion strengthens |
| pass @ θ=3.0, VIX<13 / >30 | 0.0% / 80.5% | 0.0% / 81.0% | unchanged |
| calm vs elevated pass | 18.7% / 64.2% | 18.7% / 63.9% | unchanged |
| eligible calm days/yr | 187 → 35 (−81%) | 188 → 35 (−81%) | unchanged |
| θ crossed every | 8.2 days | 8.2 days | unchanged |
| day-over-day VRP_size sd | 2.88 | 2.09 | overstated above |
The noise claim was overstated: the floored 2015 flash-crash forecasts inflated day-to-day jumps. The six live observations span 0.67, which is about a third of a typical day's move — not "under a quarter" as stated above. The conclusion (θ sits inside the noise and has no hysteresis) holds; the magnitude was exaggerated.
The HAR finding under "Moderate" — nothing checks persistence before iterating 22 steps — is fixed by the guard; the underlying cause (OLS on variance levels with a squared-return proxy) is not, and remains the reason to move to a log-HAR or range-based RV proxy through the amendment process.
Status — B3 fixed (2026-09-17)¶
max_loss now takes the wider wing. The defect was in the specification as
well as the code: 85-computations/ev-iron-condor.md prescribed
min(wing_put, wing_call) − net_credit, and the code implemented it faithfully.
Both are corrected.
- Verified against brute-force ground truth (minimum of the actual expiry P&L on
a 200,001-point grid) across a 5×5 sweep of wing widths:
references/computations/test_ev_iron_condor.py. Before the fix, exactly the 20 asymmetric cases failed and all 5 symmetric cases passed. - Only
max_losschanges, and only on asymmetric structures; every other output is bit-identical. The review case moves from −0.68 to 174.32 points; the day-1 rehearsal (symmetric) is unchanged at 89.885305. - The production engine was never affected.
ev_ladder.pyderives risk from the worst P&L over a 601-point price sweep rather than a width formula, and reports the true max loss for asymmetric condors in both orientations (tested on the selftest's own fixture). The live book's sizing was not exposed; the exposure was the attested receipts and the heat-cap input the risk function specifies.
Not addressed here: the condor validates no strike ordering, and abs() on the
wing widths would silently accept an inverted wing (long_put > short_put) as a
valid condor. Separate defect; recorded for follow-up.
Status — B2 fixed (2026-09-17)¶
The attester now re-executes the computation. After every structural check
passes (including code_version, which proves the file on disk produced the
receipt), it rebuilds the params dataclass from receipt.inputs, calls
compute_ev, and requires every recorded output to match within 1e-5.
- Tests:
references/attesters/test_ev_binding.py, 10 cases. Two forgery tests initially passed by accident — the Layer 1 flag cross-check caught an incoherent forgery — so they were rewritten as coherent forgeries (every flag and derived value consistent), which only re-execution can expose. Before the fix a coherent forgery of the negative-EV demo condor attestedverdict: pass, ev_gate: pass: the Layer 1 floor was only as strong as the attester, which trusted the numbers it gated. ev_gateis now evaluated only on re-execution-confirmed outputs; a failed receipt reportsunverified.- Historical proof: the real 2026-09-09 rehearsal receipt re-executes
bit-for-bit against its own code (
de4a658); a coherent forgery of it is refused with each fabricated value named. - All pre-existing negative tests still fail correctly, now without executing any code (re-execution is skipped on structurally unsound receipts).
Scope limit: attestation covers the four wiki computations only. The
production engine ev_ladder.py has no attested computation (REVIEW.md
section 5, item 7) — the daily round's EV numbers remain unattested.
Status — B4 fixed (2026-09-17)¶
Both legs of the term ratio now come from one FRED request (VIXCLS and
VXVCLS — VIX3M's FRED id), taking the latest date ≤ --asof on which both carry
a close. The Cboe live delayed quote is no longer fetched at all. On the ~33 US
holidays since 2012 where FRED has a VIX print but no VXV, both legs step back to
the last shared close; with no shared close the snapshot fails closed.
The bug was two failures, not one:
- Lookahead on any run after the --asof date: the 2026-09-11 record paired
the 9/10 VIX close with the 9/11 VIX3M close (18.60, fetched 9/13) — a value
unknowable at 10:31 on 9/11 — and on every historical replay it used today's
VIX3M.
- A mixed-instant ratio on live runs: the 10:31 PT cron divided yesterday's
VIX close by today's intraday VIX3M. Observable, but the ratio describes no
single moment and violates regime-definition.md ("VIX ÷ VIX3M, prior close").
Impact on the live journal — recomputed with true same-date closes:
| record | recorded ratio → term | point-in-time ratio → term |
|---|---|---|
| 2026-09-11 | 0.9591 → flat | 0.9042 → contango |
| 2026-09-15 | 0.8833 → contango | 0.8869 → contango |
| 2026-09-16 | 0.8718–0.8935 → contango | 0.8884 → contango |
| 2026-09-17 | 0.9532 → flat, TRANSITION | 0.8976 → contango, no transition |
The regime was calm-contango-vrp+ on every day from 09-09 to 09-17. The
"first TRANSITION state" of 2026-09-17 — recorded in decisions.jsonl (line 17),
log.md and REVIEW.md — was produced by the bug; it forced that day's round to
no-new-risk. The 2026-09-11 decision record's "moved from contango to flat" is
likewise false. Correction to this review's own earlier statement: the
9/17 transition was not "legitimate by luck" — it was not legitimate.
Verified: 5 tests (tools/test_regime_snapshot.py, written first; all five
reproduced the journal's exact artifacts before the fix, including the 9/11
record's vix3m 18.60 and its vix3m_asof timestamp). Live run against real
FRED: 9/16 closes 17.71 / 19.73, ratio 0.8976, calm-contango-vrp+, no
transition — where the buggy code gave calm-flat at 0.9547 on the same day.
Separate, not fixed — replay semantics. "Latest close ≤ asof" means a replay
of day D run after D's close uses D's close, while the live 10:31 PT run uses
D−1's. This affects VIX, VIX3M and the HAR forecast equally, so ratios stay
internally consistent, but replayed labels run one day ahead of live ones. Any
calibration replay must use asof = D − 1 trading day to match live behaviour.
Status — M1 and M2 fixed (2026-09-17)¶
M1, the transition window. A day is now in transition when either the raw
label differs from the confirmed one (pending confirmation) or the confirmed
label changed within transition_lookback_days business days (cooling-off),
as regime-definition.md specifies. Before, the day a change confirmed,
transition dropped to False and new risk resumed at once.
confirmation_days is honoured (it was hardcoded to 2). Consequence: a genuine
regime change now yields at least one pending day plus five business days of no
new risk. That is the documented rule; if it proves too conservative it is a
config value, changed through the amendment process.
M2, re-run idempotency. Hysteresis observations are keyed by the close
they measure (vix_date), not the run date, so a same-day re-run and a run on
a market holiday (the cron runs on weekdays, including holidays, and then sees
the prior close again) replace the observation rather than duplicating it. The
window is computed from dates, not events, so a re-run returns the same answer.
The journal skips a record identical to one already filed for that date; a
changed observation is still appended. An out-of-order replay without
--no-append no longer rewrites live state.
Records gain raw_label, confirmed_label and transition_reason (additive
v1), and the transition label no longer truncates the new label.
Verified: 9 new tests (tools/test_regime_snapshot.py, 14 total), written first
and each reproducing its defect before the fix. Simulated against a copy of
hermes' real legacy state (entries keyed by run date, the bogus 09-17 flat
entry still present) with live FRED and HAR: one observation added, no
transition, and two further re-runs changed neither state nor journal.
Remaining snapshot defect: M3, no staleness guard.
Status — M3 fixed (2026-09-17)¶
The snapshot now checks that its inputs are current, not merely present. The
FRED close (VIX and VIX3M share one date) and the HAR forecast's window_end must
each lie within max_input_age_business_days business days of --asof; a HAR
output without window_end counts as stale, since its freshness cannot be
verified. A stale input becomes a data_missing reason, so the day gets no label,
exit 2, and no hysteresis observation — the existing fail-closed path.
Threshold 2, set from evidence. Input age at cron time, 2012–2026: 1 business day on 3,678 weekday runs, 2 on 138 (the day after a holiday), 3 exactly once (2012-10-31, markets reopening after Hurricane Sandy), never more. The guard admits every normal and post-holiday run and would have failed closed only on the post-Sandy morning, where no new risk is the desired outcome anyway.
Verified: 8 tests (22 in the suite), written first; 6 failed before the fix and
the 2 gap tests (weekend, post-holiday) passed before and after by design. End to
end on real data, --asof 2026-09-23 simulates an outage and fails closed with
both inputs named — FRED close 5 business days old, HAR window_end 4.
Observed while verifying: an off-hours run can pair the prior VIX close with a
forecast that already includes today's close (Tiingo publishes before FRED).
The 10:31 PT cron uses the prior close for both, so live labels are unaffected;
forecast_window_end now makes the pairing auditable.
All four snapshot defects (B4, M1, M2, M3) are closed.
Status — M5 / REVIEW item 7: mechanism built, not yet live (2026-09-17)¶
references/attesters/ladder-binding.py attests the production engine by
re-execution: it reruns ev_ladder.build_ladder on the recorded snapshot and
trades and requires every saved ladder to match (numbers within 1e-9), recording
sha256 of the engine and all three inputs. No change to the engine or its
output. Contract: 85-computations/ev-ladder.md.
- Feasibility first: production's 2026-09-17 ladders re-execute bit-identical, 5 of 5. The engine's single wall-clock read fires only for an undated snapshot.
- On hermes, production's own ladders attest — 2026-09-17 (5/5) and
2026-09-16-evening (8/8). A coherent forgery of the real TSLA ladder is refused
with each field named (and shows the straddle's true
EV_RNis exactly 0.0, the pilot's validation signature). - 9 tests, written first, synthetic fixtures only.
Why it is not live: the round writes its trade inputs to /tmp, so every
ladder before 2026-09-16-evening is permanently unattestable — its inputs are
gone. The two surviving input files were copied into hermes-pilot/ on
2026-09-17. A proposed prompt change keeps trades in hermes-pilot/ and makes
attestation a mandatory step whose verdict the report must cite
(ops/cron/options-round-analyst.prompt.proposed.txt); it awaits owner
approval.
Scope: reproducibility only. No ladder EV gate — whether a long-vol book whose
EV_RN ≈ 0 by design needs a Layer-1 analogue is a policy question. Marks are not
validated either; the engine's own rn_warning does that.
Status — M10 fixed (2026-09-17)¶
80-system-design/hermes-cron-prompt-v2.md is rewritten. It no longer embeds prompt
text — the embedded copy is what drifted (ai-rig paths, v2.1 text while later
versions ran) — and instead explains both jobs section by section, pointing at
the canonical, live-identical text in the repository's ops/cron/. It now
documents report-builder (previously undocumented), reconstructs the version
history from the archives, and states the change procedure. Both Hermes pages
are added to the section index, where neither appeared.
Found while reconstructing the history: v2.2 and v2.3 reached production with
no governance record, and v2.3 was never archived. Also: the prompt's BOOK
CONTEXT prose is already stale (four straddles listed; the 09-17 ladders price a
fifth position), and report-builder serves the report at a public,
unauthenticated quick-tunnel URL — recorded as known issues on the page.