Summary¶
A backtest is an experiment, and an experiment repeated enough times will always find something that looks great. This concept fixes the validation rules for every strategy candidate and ML model in the system: purged walk-forward evaluation, explicit accounting for the number of trials, the deflated Sharpe ratio as the acceptance statistic, and a fixed stress set drawn from the ThetaData 2016+ history that no candidate may skip.
Overfitting and Multiple Testing¶
- Trial counting is mandatory. Every parameter tweak, feature set, and strategy variant evaluated against the same data is a trial. After enough trials, the maximum in-sample Sharpe of pure noise is high (Lopez de Prado's central result). Log trial counts per candidate family, and expect the reported Sharpe to shrink as trials accumulate[^ldp].
- Data snooping: choosing the model on the full dataset, then reporting its performance on a subset of that same data. Prevention is structural — the test window is opened once, at the end.
- Selection bias in published rules: any tradable "pattern" from a book or paper has survived implicit testing by others; treat it as a hypothesis, not evidence[^tohf-plan].
Purged, Embargoed Walk-Forward CV¶
The harness for everything — vol models, structure generators, filters:
- Walk-forward: fit/validate on a trailing window, test on the next unseen block, roll forward. Never shuffle time-series folds.
- Purging: remove training samples whose label windows overlap the test block (overlapping multi-day labels leak information across the boundary).
- Embargo: additionally drop a buffer (e.g., a few days ≈ the longest label horizon) after the training end, since serially correlated volatility makes adjacent periods informative about each other[^ldp].
- One final holdout: a never-touched recent block, evaluated once per candidate family, results recorded regardless of outcome (journaling per learning).
Deflated Sharpe Ratio¶
The deflated Sharpe ratio (Lopez de Prado, 2014) adjusts the observed Sharpe for the expected maximum Sharpe of N trials, non-normal returns (skew, kurtosis), and sample length, returning a probability that the true Sharpe exceeds zero. Acceptance rule: a candidate is eligible only if its DSR is significant after logging the family's trial count. A raw Sharpe without a deflation is not evidence.
Minimum Backtest Length and the Non-Negotiable Stress Set¶
ThetaData options data starts in 2016, giving roughly one decade of option-level history. One decade of daily data is short: at ~252 days/year, few independent volatility regimes occur, which is precisely why the evaluation set is defined by events, not days. Every candidate (model or playbook) must be evaluated across the full 2016–present window and must explicitly report behavior in all three of:
| Stress class | Example year | What it tests |
|---|---|---|
| Vol spike / crash (2018-class) | Feb 2018 "Volmageddon", Q4 2018 | Short-vol blowup risk, IV term-structure inversions, gap handling |
| Pandemic/liquidity event (2020-class) | Mar 2020 | Extreme skew, liquidity evaporation, spread widening, margin stress |
| Euphoria/low-vol grind with sharp reversals (2024-class) | Aug 2024 vol shock in a bull year | Premium-selling in complacent regimes, regime-switch speed |
This triple is non-negotiable: a strategy that only shows results averaged across regimes, or whose 2020 and 2018 behavior is unexamined, is rejected regardless of aggregate statistics. Event-class evaluation is also where the regime framework earns its keep — report metrics per regime, not only in aggregate.
References¶
- Marcos López de Prado, Advances in Financial Machine Learning, Wiley 2018 (purged k-fold CV and embargo, Ch. 7; backtesting on synthetic data, Ch. 10–12).
- Marcos López de Prado, "The Deflated Sharpe Ratio," Journal of Portfolio Management 40(5), 2014.
- Campbell Harvey & Yan Liu, "Backtesting," Journal of Portfolio Management 42(1), 2015.
- Andrew W. Lo, "The Statistics of Sharpe Ratios," Financial Analysts Journal 58(4), 2002.
- Bailey, Borwein, López de Prado & Zhu, "The Probability of Backtest Overfitting," Journal of Computational Finance, 2016.
Links¶
- Vol Forecasting Baselines and ML for Vol Prediction — the models this harness validates.
- Regimes — per-regime reporting.
- Automation Architecture — where validated models plug into execution.
- TOMIC process — the human process these gates formalize.