Skip to content

Summary

The LLM agent is a reading, writing, and process-enforcement worker, not a forecaster. It sits between unstructured inputs (news, books, journals) and the deterministic machinery (EV computations, backtests) that actually decides whether money is at risk. Its authority ends at the human approval gate; everything numeric flows through attested code.

Where LLM Agents Genuinely Fit

Role What the agent does Where it lands in the system
News triage Read headlines/filings, classify relevance (event type, tickers, vol-impact hypothesis), draft dated inbox concepts with URL, as-of time, snapshot note, and stale_after Append-only concepts in 95-inbox/index.md per the news-intake policy[^charter]
Concept retrieval Answer "what do the books say about X" by pointing into the evidence bundles and wiki, with per-claim footnotes Citations into 10-pricing-greeks, 20-volatility, 30-strategies
Candidate-structure generation Turn an inbox event + regime context into draft playbook skeletons: candidate structures, edge claims, risks, regime links — with explicit "assumptions to verify" lists Draft playbooks in 90-playbooks/index.md, status: draft
EV contract enforcement Check that every playbook states a real-world probability model, an explicit cost model, a GARCH/HAR + IV−RV baseline comparator, and links an Attested Computation; reject prose that improvises arithmetic Gate before any playbook is marked evaluable; computations run in 85-computations/index.md[^charter]
Journaling and synthesis Summarize trade journals, extract lessons, flag which rules were violated, propose edits to the wiki Journal summaries feeding learning[^tohf-learning]
Backtest plumbing Draft harness code (purged walk-forward, DSR reporting) for human review; never self-approve results Code under review; validation per backtest-discipline.md

Where LLM Agents Do NOT Fit

  • Probability estimation. Real-world probabilities for the EV contract come from named, auditable models — historical conditional frequencies, GARCH/HAR forecasts — not from a language model's intuitions. An LLM "feeling" that a 25-delta put is cheap is exactly the pattern the EV contract exists to block[^charter].
  • Price or volatility forecasting. Forecasting belongs to the baseline and ML pipeline of vol-forecasting-baselines.md and ml-for-vol-prediction.md, validated under backtest-discipline.md. An LLM's narrative plausibility is not a forecast.
  • Final risk decisions. Sizing, max-loss limits, and the discipline rules from risk management are fixed constraints; the agent may remind, never relax.
  • Unreviewed execution. The agent never places orders or mutates attested computation code without review.

The Human-in-the-Loop Approval Gate

Every path from agent output to live risk passes one gate:

  1. Agent produces a draft artifact (inbox concept, playbook, computation change) — always status: draft, always attributed by: agent/<name>.
  2. Deterministic checks run (EV contract checklist, computation execution with receipts, backtest stress set coverage from backtest-discipline.md).
  3. Human approves (verified: { by: human:yasu, ... }) or rejects with reasons; rejection reasons are journaled. Promotion of drafts to stable is exclusively a human event[^charter].

This mirrors the TOMIC philosophy in TOMIC: the process is a team product with defined review, not one person's (or one model's) reflex. The agent accelerates the reading and paperwork layers of the automation architecture; the judgment and sign-off layers remain human.

References

  • Marcos López de Prado, Advances in Financial Machine Learning, Wiley 2018 (why narratives are not backtests; code review discipline).
  • Open Knowledge Format v0.2 specification (../OKF_SPEC_v0.2.md) — actor conventions, EV contract, news-intake policy.
  • Chen & Sebastian, The Option Trader's Hedge Fund — TOMIC process, infrastructure, and learning chapters (see evidence bundle topics cited above).

Links