Summary¶
Machine learning enters the system only after the GARCH/HAR baselines of vol-forecasting-baselines.md exist and an honest out-of-sample harness (backtest-discipline.md) is running. The promising surface is not raw return prediction but conditioning the IV−RV spread and the vol forecast on information the classic models ignore: the shape of the IV surface, options flow, and market microstructure.
Feature Ideas from the IV Surface and Market Data¶
The books supply the intuition for most useful features — IV is set by supply and demand, and its shape (ATM level, skew, term structure) moves in patterns traders already read[^tohf-vol][^hull-smiles].
| Feature group | Examples | Rationale |
|---|---|---|
| IV surface shape | ATM IV level and z-score; put–call skew (e.g., 25Δ put − 25Δ call); term slope (front vs back month IV) | Skew and term structure summarize market fear/premium supply[^tohf-vol] |
| Variance risk premium | IV (30d) − trailing RV; IV − GARCH/HAR forecast | The tradable spread; VRP level is a known conditioning variable |
| Realized-vol state | lagged RV (1d/1w/1m), Parkinson/GK ranges, RV-of-RV | HAR inputs; let the model learn nonlinearity in them |
| Flow and positioning | option volume, open interest changes, put/call volume ratio | Demand pressure moves IV[^tohf-vol] |
| Microstructure | bid–ask spreads, TOMIC-style "IV rank/percentile" analogues, trend of IV over days1 | Chosen inputs for entries in the TOMIC process |
| Calendar/event | earnings and macro calendar dummies, days-to-event | Term-structure distortions around events[^tohf-vol] |
Target choices: forecast future RV (regression), forecast the sign/magnitude of future IV change, or forecast the future VRP — the last is closest to tradable edge.
Model Classes¶
| Class | Candidates | When appropriate |
|---|---|---|
| Linear, regularized | ridge / lasso on standardized features | Transparent benchmark upgrade over HAR; resists overfitting given small effective samples |
| Tree ensembles | random forest, gradient-boosted machines (XGBoost/LightGBM) | Capture interactions (e.g., skew × term slope) with modest data; must be heavily regularized |
| Sequence models | LSTM / GRU / temporal CNNs on rolling windows | Only if the pipeline is mature; high overfit risk, low interpretability |
| Anything else | — | Not before the above are beaten honestly |
Non-negotiables for every class: purged walk-forward cross-validation with embargo, no leakage of same-day close information into same-day features, and comparison against GARCH(1,1) and HAR-RV on identical windows.
Realistic Gains¶
Expect modest improvements: a few percent of RMSE, or slightly better tail calibration — not a step change. GARCH and HAR already encode the two dominant facts (clustering and multi-scale persistence); the marginal information in surface shape and flow is real but small and partly arbitraged. A model earns its place only if the improvement survives the stress set and multiple-testing corrections of backtest-discipline.md, and if it translates into better EV versus the baseline comparator required by the EV contract (source.md). Prioritize the VRP/IV-change targets, where small forecast gains convert most directly into premium-selling decisions in the 80-system-design automation.
Pitfalls¶
- Nonstationarity: relationships between surface features and future vol shift with regime; retrain on rolling windows and evaluate per-regime (60-regimes/index.md).
- Lookahead bias: using a feature (e.g., same-day IV close) that wasn't knowable at decision time; ThetaData timestamps must be respected.
- Small effective sample: daily vol data is highly autocorrelated — thousands of rows may contain only a handful of independent vol regimes; this inflates apparent skill.
- Overlapping targets: multi-horizon targets are serially correlated; standard CV folds overstate accuracy without purging.
- Target leakage via labels: labels built from revised/cleaned data the model could not have seen live.
References¶
- Lopez de Prado, Advances in Financial Machine Learning, Wiley 2018 (chapters 5–8: purged CV, feature importance, backtest overfitting).
- Gu, Kelly & Xiu, "Empirical Asset Pricing via Machine Learning," Review of Financial Studies 33(5), 2020.
- Corsi, "A Simple Approximate Long-Memory Model of Realized Volatility," Journal of Financial Econometrics 7(2), 2009.
- Bollerslev, Tauchen & Zhou, "Expected Stock Returns and Variance Risk Premia," Review of Financial Studies 22(11), 2009.
- Christoffersen, P. F., & Diebold, F. X. (2000). How Relevant is Volatility Forecasting for Financial Risk Management? Review of Economics and Statistics, 82(1), 1–11.
Links¶
- Vol Forecasting Baselines — the bar to beat.
- Backtest Discipline — validation harness.
- Regimes — regime conditioning of features and evaluation.
- Automation Architecture — where forecasts feed the system.
-
Chen & Sebastian, The Option Trader's Hedge Fund — infrastructure topic (
sources[]: tohf-infra). ↩