Skip to content

Review findings — 2026-09-17

Review of the framework as it runs in production. Findings are open unless marked otherwise; this repository's initial commit is the "before" baseline, so each fix lands as its own reviewable diff.

Every claim below was reproduced by running the code, not read off the source.


Blockers

B1 — The EV gate is vacuous, and inverted relative to the VRP thesis (Layer 1 FIXED 2026-09-17; Layers 2–4 proposed)

edge_vs_rv = ev_net(σ=garch) − ev_net(σ=rv) (references/computations/ev_iron_condor.py:206). Both terms price the same structure under two vol assumptions. A short condor's EV is monotonically decreasing in σ, so the gate reduces to one bit:

edge_vs_rv > 0 ⟺ garch_forecast < rv_window

Reproduced:

IV RV HAR ev_net edge_vs_rv Gate
0.30 0.12 0.15 +26.96 −11.82 FAIL
0.12 0.20 0.15 −16.15 +16.52 PASS

Row 1: an 18-vol-point VRP with +27 EV — declined. Row 2: selling vol below both realized and forecast, negative EV — approved. The gate passes money-losers and rejects the best VRP trades whenever the forecast sits above trailing realized, which is the normal state coming out of a calm stretch.

REVIEW.md §3 reads the day-1 decline as "the edge test rejects trades naive premium-selling would take." What actually happened was 0.11 > 0.0952.

The root cause is narrower than the comparator's direction: no gate anywhere tested ev_net > 0. That absence is what admitted row 2.

Status: resolved by Amendment 2026-09-17/01, which retires edge_vs_rv as a gate (retaining it as a diagnostic) and replaces it with a layered test:

Layer Test Status
1 ev_net > 0, net of costs in force 2026-09-17
2 iv − forecast_rv ≥ θ pending θ calibration
3 edge survives the forecast's standard error pending se_forecast
4 conditional EV beats the unconditional archetype pending the backtest harness

Layer 1 is implemented: all four computations emit gate_ev_net_positive, and the attester reports a separate ev_gate derived from the mandatory ev_net — backward compatible with pre-amendment receipts — exiting non-zero when either check fails. Verified 5/5 on the cases above; tampered-flag case negative-tested.

The dual-measure alternative EV_PHYS − EV_RN was tested and rejected: EV_RN ≈ −costs, not ≈ 0, so subtracting it restores the cost drag and passes losing trades (5 of 5 in the thin-edge regime).

B2 — The attester does not attest

references/attesters/ev-binding.py validates key names, types, path binding and sha256(computation file). It never re-runs the computation. Taking the real day-1 receipt and flipping ev_net 12.73 → 9999, edge_vs_rv −5.75 → +9998, inputs untouched:

{"verdict": "pass", "reasons": []}   exit=0

Charter principle 2 is "No attestation, no proposal"; the stated failure mode it prevents is "LLM arithmetic, non-reproducible numbers." A fabricated receipt passes.

Fix: import the module, call compute_ev(**receipt["inputs"]), compare with receipt["outputs"] within tolerance. Roughly ten lines.

Related: code_version = sha256(file) means any edit invalidates every historical receipt with no recovery path. Retain versioned source, not just its hash.

B3 — max_loss reports a risk-free position that can lose $17k

ev_iron_condor.py:194 takes min() of the two wings. With put wing 200 and call wing 25:

reported max_loss:  -0.68     <- negative: "no maximum loss"
true max_loss:     174.32     <- $17,432 per contract

max_loss feeds position sizing and the 20% heat cap. ev_butterfly.py rejects asymmetric wings; the condor accepts and misreports them.

Fix: max(), and subtract costs.

B4 — Lookahead in the live regime journal

tools/regime_snapshot.py:64 — vix3m_from_cboe() takes no --asof and fetches the current delayed quote. Visible in the journal:

date: 2026-09-11   vix_date: 2026-09-10   vix3m_asof: 2026-09-13 01:52:51

The 9/11 term ratio was computed from 9/13 data. That ratio was 0.9591 — 1% from flipping flat→contango and changing the label. regime-definition.md calls ex-ante computability non-negotiable. Any point-in-time replay used to calibrate θ inherits this.

Also d["data"]["close"] or d["data"]["current_price"] mixes prior-close and intraday depending on run time.


Majors

M1 — The framework's most conservative rule is not implemented. transition_lookback_days: 5 appears in tools/regime_config.json and nowhere in the code. The docs define a transition as "the confirmed label changed within the last N days" — a cooling-off window. The code defines it as raw ≠ confirmed now, so the day a flip confirms, transition goes False and new risk resumes immediately. confirmation_days is equally decorative: line 147 compares it to len(recent) but only tests recent[-1] == recent[-2], hardcoding 2.

M2 — Same-day re-runs can falsely confirm a regime label. A duplicate append makes recent[-1] == recent[-2] trivially true. Live evidence: regime-state.json lists 2026-09-16 three times. In this instance 9/15 and 9/16 were both calm-contango-vrp+, so the 9/17 transition was legitimate — right answer, unsound mechanism. The log records this as "minor idempotency"; it gates Gate 1 eligibility.

M3 — Fail-closed catches missing data, never stale data. vix_from_fred scans back 45 days and returns the newest value found, reporting data_missing: false. A FRED outage yields a three-week-old VIX indistinguishable from a fresh one. Same for vix3m_ts and har["window_end"]. Needs a max-age assertion per input.

M4 — outcomes.jsonl does not exist on either host. The learning function exists to close forecast → reality, and the concordance quorum needs ≥3 concordant outcome events. No file, no writer, Hermes loop stage 10 undone. The amendment process is unreachable by its primary path.

M5 — The deployed EV engine is unattested. ev_ladder.py won the consolidation but has no AC concept and emits ladder JSON the attester cannot parse, while hermes-cron-prompt-v2.md is status: stable and applied. The live daily round contradicts charter principle 2 continuously. Should be a recorded exception with an end date, not implicit.

M6 — Journal fork (healed in this repo's initial commit). hermes held regime.jsonl (9 records) and no decisions.jsonl; the workstation held decisions.jsonl (17 records) and a 2-record regime.jsonl. Neither was a superset. Merged here; single-writer rule now documented in README.md.

M7 — Six of seven framework documents are absent from the production host. hermes has only decision-function.md (an older revision, missing the regime_config.json canonical reference). risk-function.md — the never-rules, the frozen v0 risk parameters, the standing tail rules — is not present where the agent runs. log.md 2026-09-16 claims "canonical wiki has all 88-framework docs on both machines"; that claim is false for hermes. Fixed by deploy.sh.

M8 — No actor field in the decision log. Records assert "owner approved" inside agent-written free text. For the one inviolable rule — the human gate — there is no machine-checkable distinction between a human approved and the agent wrote that a human approved. The OKF spec has an actor convention the decision schema does not use.

M9 — 82 of 285 relative provenance links are broken (one ../ short); four files over-corrected to ../../../../. log.md records fixing these in sources[] frontmatter; body footnotes were missed. Provenance is the wiki's headline feature.

M10 — Documented prod config does not match deployed prod config. hermes-cron-prompt-v2.md and the archived cron-prompt-v2.*-applied.txt copies all reference /home/ysakakibara/...; the live jobs.json uses /home/hermes/... and is 7003 chars against v2.2's archived 6744. A second job, report-builder, is undocumented. Both live prompts are now exported to ops/cron/.


Moderate

  • HAR fit quality. Squared daily close-to-close returns are a very noisy RV proxy — errors-in-variables will attenuate b_daily; OLS on variance levels lets a few crash days dominate. Corsi fits intraday RV, usually in logs. Tiingo already returns OHLC, so a Garman–Klass estimator is ~5× more efficient for free. Nothing checks b_d + b_w + b_m < 1 before iterating 22 steps.
  • Two incompatible VRP definitions. Snapshot: VIX − HAR_forecast. EV code: iv − rv_window (trailing). Both called VRP; θ is calibrated against the first.
  • θ is source-conditional. VIX is a variance-swap strike, not ATM IV30. When ThetaData IV30 lands (V1b) the VRP level shifts systematically by ~1–2 vol points and any θ calibrated now becomes invalid.
  • Default slippage_pct = 0.15 — $527 on a $3,511 SPX condor credit — contradicts the charter's own "zero commissions, penny-wide SPX" row.
  • EV models hold-to-expiry while the pilot's headline lesson is "the exit path IS the edge" (SPY 765 straddle, EV_PHYS −21.9%). The attested number does not describe the trade being proposed.
  • Drift drops the dividend yield (μ = r + ERP, ~1.2%/yr high on a price index) and e_value is an undiscounted terminal expectation added to a PV credit. Both small, both systematic.
  • Tiingo key passed in the URL query string (tools/har_forecast.py) — use the Authorization: Token header.
  • REVIEW.md redefines the calibration gate. Open decision 2 says calibration "awaits sufficient VRP_size series," citing three live prints. The amendment's own plan requires a 2016→present ex-ante series with purged walk-forward and the 2018/2020/2024 stress set — runnable today, since har_forecast.py --asof already loops. The real blocker is the absent backtest harness.
  • The decision log is ~90% self-administration. Of 17 records, 2 concern trades; the rest are cron prompts, env hardening, ledger corrections, migration verification.

What holds up

  • 60-regimes/amendment-2026-09-09-vrp-size-threshold.md disqualifies its own supporting evidence, downgrades its quorum claim, and pre-commits to "or the honest conclusion that no θ works." Rare discipline.
  • Fail-closed as a first-class concept; missing data as a state, not an assumption.
  • regime-detection-methods.md — filtered vs smoothed, published detection-delay distributions, ±20% threshold perturbation.
  • Forced resolution, versioned rules, no mid-session amendment: the right structural answer to LLM drift.
  • The research brief was allowed to undercut the project's own premise (VRP alphas ≈ 0 post-2012) and drive a redesign rather than a defence.

Suggested order

  1. B2 — make the attester re-execute the computation
  2. B3 — max() of wings
  3. B4 + M3 — thread --asof to VIX3M; max-age guards on all inputs
  4. M1 + M2 — implement transition_lookback_days, honour confirmation_days, key recent by date so re-runs are idempotent
  5. M4 — create outcomes.jsonl and its writer
  6. M9 — fix the links (one sed pass)
  7. ~~B1 — redesign the edge test~~ — Layer 1 done; Layers 2–4 await θ calibration, se_forecast, and the backtest harness

One piece of work unblocks the most: the calibration backtest harness gates both the θ calibration (Amendment 2026-09-09/01) and Layer 4 of the edge test. It is the highest-leverage item on this list.

1–6 are roughly a week. None of the six framework concepts should be promoted to stable before they land: the drafts currently describe a more rigorous system than the one running.


Appendix — Empirical Basis for the Amendment 2026-09-09/01 Review

Added 2026-09-17. Reproduce with tools/vrp_series.py (steps 1–2 of the amendment's own calibration plan), then the analysis below. Inputs are the same proxies the production tool uses — VIX as IV30, SPY close-to-close RV — so these figures describe what this rule does with this pipeline. Effect ranking is robust; exact percentages inherit the proxies' limitations.

Series: 3,396 ex-ante daily observations, 2013-03-18 → 2026-09-16. HAR-RV refit on trailing 756 observations at each date, data ≤ t only.

θ = 3.0 is substantially a VIX-level filter

corr(VIX, VRP_size) = +0.476; slope +0.288 vol points of VRP per VIX point.

VIX band days pass @ θ=3.0 mean VRP
<13 675 0.0% 0.09
13–16 993 14.0% 1.34
16–20 851 39.0% 1.93
20–25 484 52.7% 3.38
25–30 235 77.4% 5.69
>30 149 80.5% 7.39

Against the mapping it serves: 18.7% of calm days (VIX<20, premium selling eligible) versus 64.2% of elevated days (premium selling forbidden). θ is most permissive where the mapping forbids trading and most restrictive where it permits it. Over 675 days at VIX<13, not one day would have been eligible.

Undisclosed consequence: eligible calm days fall from 187/yr to 35/yr, an 81% reduction.

θ sits inside the forecast noise

Day-over-day VRP_size change: sd 2.88 vol points. The threshold is crossed every 8.2 days; 22.4% of all days fall within ±1.0 of it. The six live observations (2.68 / 2.78 / 2.84 / 2.92 / 3.29 / 3.35) span 0.67 — under a quarter of one day's typical move. They sample noise around θ rather than converging on it. No hysteresis is applied to this axis, unlike the VIX and term axes, contrary to regime-definition.md.

The calibration cannot produce the number it asks for

A 45-DTE trade occupies ~31 trading days, so 2013–2026 yields ~109 independent trades: ~17 in bucket 0–1, ~32 in 1–3, ~21 in 3–6, ~11 in >6. The mandatory stress set is ~1 independent trade per episode, 3 total — a sanity check, not a statistical test. Selecting "the smallest bucket boundary at which net edge turns positive" across those buckets is threshold-mining; the plan requires a multiple-testing control but names none.

The forecast breaks exactly where the stress set lives

All 9 implausible forecasts in 13 years fall inside stress windows:

2020-03-10  124%     2020-03-13  559%     2020-03-24  165%
2020-03-11  116%     2020-03-16 1008%     2025-04-09  133%
2020-03-12  466%     2020-03-17  461%     2025-04-10  119%

The mandatory stress-set validation would run on inputs like a 1008% annualized vol forecast. Missing stationarity check on the iterated HAR path.

Scale-invariance helps but does not close it

VRP_ratio = VIX/forecast − 1 at a matched 30.4% pass rate drops corr(VIX, ·) from +0.476 to +0.292 and flattens the bands (1.2 / 24.2 / 38.0 / 44.4 / 60.9 / 67.8%) — but they remain monotone. VRP is genuinely regime-dependent, so no single global θ is correct; θ must be conditioned on the regime band. The calibration plan's strata are on VRP_size buckets when the variable requiring stratification is the regime itself.

Correction — re-run under the HAR validity guard (2026-09-17, later)

The figures above were computed on a hand-filtered series (days with a forecast above 100% or below 2% dropped). tools/har_forecast.py now carries a principled validity guard that refuses 20 days — explosive fits (companion spectral radius ≥ 1) and stable-but-sign-flipped fits that predict negative variance, which the hand filter kept (e.g. 2015-08-31 at 8.3%). Re-run on the guarded series (3,376 days, 20 excluded):

Figure hand-filtered (above) guarded
corr(VIX, VRP_size) +0.476 +0.498 conclusion strengthens
pass @ θ=3.0, VIX<13 / >30 0.0% / 80.5% 0.0% / 81.0% unchanged
calm vs elevated pass 18.7% / 64.2% 18.7% / 63.9% unchanged
eligible calm days/yr 187 → 35 (−81%) 188 → 35 (−81%) unchanged
θ crossed every 8.2 days 8.2 days unchanged
day-over-day VRP_size sd 2.88 2.09 overstated above

The noise claim was overstated: the floored 2015 flash-crash forecasts inflated day-to-day jumps. The six live observations span 0.67, which is about a third of a typical day's move — not "under a quarter" as stated above. The conclusion (θ sits inside the noise and has no hysteresis) holds; the magnitude was exaggerated.

The HAR finding under "Moderate" — nothing checks persistence before iterating 22 steps — is fixed by the guard; the underlying cause (OLS on variance levels with a squared-return proxy) is not, and remains the reason to move to a log-HAR or range-based RV proxy through the amendment process.

Status — B3 fixed (2026-09-17)

max_loss now takes the wider wing. The defect was in the specification as well as the code: 85-computations/ev-iron-condor.md prescribed min(wing_put, wing_call) − net_credit, and the code implemented it faithfully. Both are corrected.

  • Verified against brute-force ground truth (minimum of the actual expiry P&L on a 200,001-point grid) across a 5×5 sweep of wing widths: references/computations/test_ev_iron_condor.py. Before the fix, exactly the 20 asymmetric cases failed and all 5 symmetric cases passed.
  • Only max_loss changes, and only on asymmetric structures; every other output is bit-identical. The review case moves from −0.68 to 174.32 points; the day-1 rehearsal (symmetric) is unchanged at 89.885305.
  • The production engine was never affected. ev_ladder.py derives risk from the worst P&L over a 601-point price sweep rather than a width formula, and reports the true max loss for asymmetric condors in both orientations (tested on the selftest's own fixture). The live book's sizing was not exposed; the exposure was the attested receipts and the heat-cap input the risk function specifies.

Not addressed here: the condor validates no strike ordering, and abs() on the wing widths would silently accept an inverted wing (long_put > short_put) as a valid condor. Separate defect; recorded for follow-up.

Status — B2 fixed (2026-09-17)

The attester now re-executes the computation. After every structural check passes (including code_version, which proves the file on disk produced the receipt), it rebuilds the params dataclass from receipt.inputs, calls compute_ev, and requires every recorded output to match within 1e-5.

  • Tests: references/attesters/test_ev_binding.py, 10 cases. Two forgery tests initially passed by accident — the Layer 1 flag cross-check caught an incoherent forgery — so they were rewritten as coherent forgeries (every flag and derived value consistent), which only re-execution can expose. Before the fix a coherent forgery of the negative-EV demo condor attested verdict: pass, ev_gate: pass: the Layer 1 floor was only as strong as the attester, which trusted the numbers it gated.
  • ev_gate is now evaluated only on re-execution-confirmed outputs; a failed receipt reports unverified.
  • Historical proof: the real 2026-09-09 rehearsal receipt re-executes bit-for-bit against its own code (de4a658); a coherent forgery of it is refused with each fabricated value named.
  • All pre-existing negative tests still fail correctly, now without executing any code (re-execution is skipped on structurally unsound receipts).

Scope limit: attestation covers the four wiki computations only. The production engine ev_ladder.py has no attested computation (REVIEW.md section 5, item 7) — the daily round's EV numbers remain unattested.

Status — B4 fixed (2026-09-17)

Both legs of the term ratio now come from one FRED request (VIXCLS and VXVCLS — VIX3M's FRED id), taking the latest date ≤ --asof on which both carry a close. The Cboe live delayed quote is no longer fetched at all. On the ~33 US holidays since 2012 where FRED has a VIX print but no VXV, both legs step back to the last shared close; with no shared close the snapshot fails closed.

The bug was two failures, not one: - Lookahead on any run after the --asof date: the 2026-09-11 record paired the 9/10 VIX close with the 9/11 VIX3M close (18.60, fetched 9/13) — a value unknowable at 10:31 on 9/11 — and on every historical replay it used today's VIX3M. - A mixed-instant ratio on live runs: the 10:31 PT cron divided yesterday's VIX close by today's intraday VIX3M. Observable, but the ratio describes no single moment and violates regime-definition.md ("VIX ÷ VIX3M, prior close").

Impact on the live journal — recomputed with true same-date closes:

record recorded ratio → term point-in-time ratio → term
2026-09-11 0.9591 → flat 0.9042 → contango
2026-09-15 0.8833 → contango 0.8869 → contango
2026-09-16 0.8718–0.8935 → contango 0.8884 → contango
2026-09-17 0.9532 → flat, TRANSITION 0.8976 → contango, no transition

The regime was calm-contango-vrp+ on every day from 09-09 to 09-17. The "first TRANSITION state" of 2026-09-17 — recorded in decisions.jsonl (line 17), log.md and REVIEW.md — was produced by the bug; it forced that day's round to no-new-risk. The 2026-09-11 decision record's "moved from contango to flat" is likewise false. Correction to this review's own earlier statement: the 9/17 transition was not "legitimate by luck" — it was not legitimate.

Verified: 5 tests (tools/test_regime_snapshot.py, written first; all five reproduced the journal's exact artifacts before the fix, including the 9/11 record's vix3m 18.60 and its vix3m_asof timestamp). Live run against real FRED: 9/16 closes 17.71 / 19.73, ratio 0.8976, calm-contango-vrp+, no transition — where the buggy code gave calm-flat at 0.9547 on the same day.

Separate, not fixed — replay semantics. "Latest close ≤ asof" means a replay of day D run after D's close uses D's close, while the live 10:31 PT run uses D−1's. This affects VIX, VIX3M and the HAR forecast equally, so ratios stay internally consistent, but replayed labels run one day ahead of live ones. Any calibration replay must use asof = D − 1 trading day to match live behaviour.

Status — M1 and M2 fixed (2026-09-17)

M1, the transition window. A day is now in transition when either the raw label differs from the confirmed one (pending confirmation) or the confirmed label changed within transition_lookback_days business days (cooling-off), as regime-definition.md specifies. Before, the day a change confirmed, transition dropped to False and new risk resumed at once. confirmation_days is honoured (it was hardcoded to 2). Consequence: a genuine regime change now yields at least one pending day plus five business days of no new risk. That is the documented rule; if it proves too conservative it is a config value, changed through the amendment process.

M2, re-run idempotency. Hysteresis observations are keyed by the close they measure (vix_date), not the run date, so a same-day re-run and a run on a market holiday (the cron runs on weekdays, including holidays, and then sees the prior close again) replace the observation rather than duplicating it. The window is computed from dates, not events, so a re-run returns the same answer. The journal skips a record identical to one already filed for that date; a changed observation is still appended. An out-of-order replay without --no-append no longer rewrites live state.

Records gain raw_label, confirmed_label and transition_reason (additive v1), and the transition label no longer truncates the new label.

Verified: 9 new tests (tools/test_regime_snapshot.py, 14 total), written first and each reproducing its defect before the fix. Simulated against a copy of hermes' real legacy state (entries keyed by run date, the bogus 09-17 flat entry still present) with live FRED and HAR: one observation added, no transition, and two further re-runs changed neither state nor journal.

Remaining snapshot defect: M3, no staleness guard.

Status — M3 fixed (2026-09-17)

The snapshot now checks that its inputs are current, not merely present. The FRED close (VIX and VIX3M share one date) and the HAR forecast's window_end must each lie within max_input_age_business_days business days of --asof; a HAR output without window_end counts as stale, since its freshness cannot be verified. A stale input becomes a data_missing reason, so the day gets no label, exit 2, and no hysteresis observation — the existing fail-closed path.

Threshold 2, set from evidence. Input age at cron time, 2012–2026: 1 business day on 3,678 weekday runs, 2 on 138 (the day after a holiday), 3 exactly once (2012-10-31, markets reopening after Hurricane Sandy), never more. The guard admits every normal and post-holiday run and would have failed closed only on the post-Sandy morning, where no new risk is the desired outcome anyway.

Verified: 8 tests (22 in the suite), written first; 6 failed before the fix and the 2 gap tests (weekend, post-holiday) passed before and after by design. End to end on real data, --asof 2026-09-23 simulates an outage and fails closed with both inputs named — FRED close 5 business days old, HAR window_end 4.

Observed while verifying: an off-hours run can pair the prior VIX close with a forecast that already includes today's close (Tiingo publishes before FRED). The 10:31 PT cron uses the prior close for both, so live labels are unaffected; forecast_window_end now makes the pairing auditable.

All four snapshot defects (B4, M1, M2, M3) are closed.

Status — M5 / REVIEW item 7: mechanism built, not yet live (2026-09-17)

references/attesters/ladder-binding.py attests the production engine by re-execution: it reruns ev_ladder.build_ladder on the recorded snapshot and trades and requires every saved ladder to match (numbers within 1e-9), recording sha256 of the engine and all three inputs. No change to the engine or its output. Contract: 85-computations/ev-ladder.md.

  • Feasibility first: production's 2026-09-17 ladders re-execute bit-identical, 5 of 5. The engine's single wall-clock read fires only for an undated snapshot.
  • On hermes, production's own ladders attest — 2026-09-17 (5/5) and 2026-09-16-evening (8/8). A coherent forgery of the real TSLA ladder is refused with each field named (and shows the straddle's true EV_RN is exactly 0.0, the pilot's validation signature).
  • 9 tests, written first, synthetic fixtures only.

Why it is not live: the round writes its trade inputs to /tmp, so every ladder before 2026-09-16-evening is permanently unattestable — its inputs are gone. The two surviving input files were copied into hermes-pilot/ on 2026-09-17. A proposed prompt change keeps trades in hermes-pilot/ and makes attestation a mandatory step whose verdict the report must cite (ops/cron/options-round-analyst.prompt.proposed.txt); it awaits owner approval.

Scope: reproducibility only. No ladder EV gate — whether a long-vol book whose EV_RN ≈ 0 by design needs a Layer-1 analogue is a policy question. Marks are not validated either; the engine's own rn_warning does that.

Status — M10 fixed (2026-09-17)

80-system-design/hermes-cron-prompt-v2.md is rewritten. It no longer embeds prompt text — the embedded copy is what drifted (ai-rig paths, v2.1 text while later versions ran) — and instead explains both jobs section by section, pointing at the canonical, live-identical text in the repository's ops/cron/. It now documents report-builder (previously undocumented), reconstructs the version history from the archives, and states the change procedure. Both Hermes pages are added to the section index, where neither appeared.

Found while reconstructing the history: v2.2 and v2.3 reached production with no governance record, and v2.3 was never archived. Also: the prompt's BOOK CONTEXT prose is already stale (four straddles listed; the 09-17 ladders price a fifth position), and report-builder serves the report at a public, unauthenticated quick-tunnel URL — recorded as known issues on the page.