Skip to content

Summary

Machine learning enters the system only after the GARCH/HAR baselines of vol-forecasting-baselines.md exist and an honest out-of-sample harness (backtest-discipline.md) is running. The promising surface is not raw return prediction but conditioning the IV−RV spread and the vol forecast on information the classic models ignore: the shape of the IV surface, options flow, and market microstructure.

Feature Ideas from the IV Surface and Market Data

The books supply the intuition for most useful features — IV is set by supply and demand, and its shape (ATM level, skew, term structure) moves in patterns traders already read[^tohf-vol][^hull-smiles].

Feature group Examples Rationale
IV surface shape ATM IV level and z-score; put–call skew (e.g., 25Δ put − 25Δ call); term slope (front vs back month IV) Skew and term structure summarize market fear/premium supply[^tohf-vol]
Variance risk premium IV (30d) − trailing RV; IV − GARCH/HAR forecast The tradable spread; VRP level is a known conditioning variable
Realized-vol state lagged RV (1d/1w/1m), Parkinson/GK ranges, RV-of-RV HAR inputs; let the model learn nonlinearity in them
Flow and positioning option volume, open interest changes, put/call volume ratio Demand pressure moves IV[^tohf-vol]
Microstructure bid–ask spreads, TOMIC-style "IV rank/percentile" analogues, trend of IV over days1 Chosen inputs for entries in the TOMIC process
Calendar/event earnings and macro calendar dummies, days-to-event Term-structure distortions around events[^tohf-vol]

Target choices: forecast future RV (regression), forecast the sign/magnitude of future IV change, or forecast the future VRP — the last is closest to tradable edge.

Model Classes

Class Candidates When appropriate
Linear, regularized ridge / lasso on standardized features Transparent benchmark upgrade over HAR; resists overfitting given small effective samples
Tree ensembles random forest, gradient-boosted machines (XGBoost/LightGBM) Capture interactions (e.g., skew × term slope) with modest data; must be heavily regularized
Sequence models LSTM / GRU / temporal CNNs on rolling windows Only if the pipeline is mature; high overfit risk, low interpretability
Anything else — Not before the above are beaten honestly

Non-negotiables for every class: purged walk-forward cross-validation with embargo, no leakage of same-day close information into same-day features, and comparison against GARCH(1,1) and HAR-RV on identical windows.

Realistic Gains

Expect modest improvements: a few percent of RMSE, or slightly better tail calibration — not a step change. GARCH and HAR already encode the two dominant facts (clustering and multi-scale persistence); the marginal information in surface shape and flow is real but small and partly arbitraged. A model earns its place only if the improvement survives the stress set and multiple-testing corrections of backtest-discipline.md, and if it translates into better EV versus the baseline comparator required by the EV contract (source.md). Prioritize the VRP/IV-change targets, where small forecast gains convert most directly into premium-selling decisions in the 80-system-design automation.

Pitfalls

  • Nonstationarity: relationships between surface features and future vol shift with regime; retrain on rolling windows and evaluate per-regime (60-regimes/index.md).
  • Lookahead bias: using a feature (e.g., same-day IV close) that wasn't knowable at decision time; ThetaData timestamps must be respected.
  • Small effective sample: daily vol data is highly autocorrelated — thousands of rows may contain only a handful of independent vol regimes; this inflates apparent skill.
  • Overlapping targets: multi-horizon targets are serially correlated; standard CV folds overstate accuracy without purging.
  • Target leakage via labels: labels built from revised/cleaned data the model could not have seen live.

References

  • Lopez de Prado, Advances in Financial Machine Learning, Wiley 2018 (chapters 5–8: purged CV, feature importance, backtest overfitting).
  • Gu, Kelly & Xiu, "Empirical Asset Pricing via Machine Learning," Review of Financial Studies 33(5), 2020.
  • Corsi, "A Simple Approximate Long-Memory Model of Realized Volatility," Journal of Financial Econometrics 7(2), 2009.
  • Bollerslev, Tauchen & Zhou, "Expected Stock Returns and Variance Risk Premia," Review of Financial Studies 22(11), 2009.
  • Christoffersen, P. F., & Diebold, F. X. (2000). How Relevant is Volatility Forecasting for Financial Risk Management? Review of Economics and Statistics, 82(1), 1–11.

Links


  1. Chen & Sebastian, The Option Trader's Hedge Fund — infrastructure topic (sources[]: tohf-infra). ↩