Skip to content

Summary

Given the state variables defined in Regime Definition, a system still has to choose how labels are produced: hand-set rules, statistical state models (HMMs), or change-point detection. The decisive comparison is not in-sample accuracy but detection latency and ex-ante validity — every method confirms a new regime only after it has begun, and the label must be generated from data available at the time the label is used1. This concept catalogues the three families honestly, including how long each typically takes to flip a label.

Rule-Based Bands

Hand-set thresholds on the state variables (VIX band, term slope, VRP sign) with confirmation rules and hysteresis.

  • How: fixed bands or rolling-percentile bands on prior-close data; enter at threshold X, exit at threshold X±buffer; require N consecutive days.
  • Latency: typically 1–5 trading days for a real regime shift (confirmation lag is a deliberate price paid for stability against flicker). Fast breakouts (Volmageddon-class) can flip the label in 1 day if the confirmation rule is permissive, but permissive rules also mislabel spikes that mean-revert within days.
  • Strengths: transparent, trivially point-in-time auditable, no fit to overfit, no lookahead, cheap. Every book-derived condition ("calm", "typhoon")5 maps directly onto a band.
  • Weaknesses: bands are static; multivariate interactions (high VIX but steep contango) require manual case rules; thresholds are judgment, so document them.

Hidden Markov Models

Fit a small (2–4 state) HMM to returns and/or vol features; the regime is the filtered most-likely state1.

  • How: features = daily index returns, realized vol, VIX level and slope. Fit by EM on trailing data; at decision time use the filtered probability P(stateₜ | data ≤ t) — never the smoothed probability P(stateₜ | full sample), which is lookahead.
  • Latency: an HMM's filtered state probability reacts within ~1–3 days to a sharp shift, but EM fits on trailing windows mean the model itself is refit with lag; a state discovered in-sample may not exist in the refit until weeks later. Persistent low-vol/high-vol states detect reliably; short V-shaped events are mostly missed until after they resolve.
  • Strengths: probabilistic output (use P>0.7 gates rather than hard labels), captures correlation shifts, principled multivariate treatment.
  • Weaknesses: label switching across refits (state 1 vs state 2 are arbitrary), fit instability, and the classic trap: in-sample decoded regimes look beautifully separated and are worthless as a trading signal.

Change-Point Detection

CUSUM and Bayesian Online Changepoint Detection (BOCPD) flag when the data-generating process shifted, rather than classifying states32.

  • How: CUSUM accumulates deviations from a running baseline and alarms when the cumulative statistic crosses a design threshold (tuned via average run length). BOCPD maintains a run-length posterior, online, one observation at a time.
  • Latency: CUSUM's latency is a tunable tradeoff — the faster the detection, the higher the false-alarm rate; expect several days of accumulation for vol-regime shifts that build over days, near-immediate response to jumps. BOCPD reacts in ~1–2 days to clear level shifts with an explicit probability of change.
  • Strengths: honest about the transition moment; the transition signal itself is the output — which maps directly to the "no new risk during transitions" rule in Regime-Dependent Delta Exposure.
  • Weaknesses: detection only, not classification (you still need bands or an HMM to say which regime); CUSUM threshold tuning is subtle.

Comparison

Method Typical detection delay Lookahead traps Output Practical role
Rule bands 1–5 days (confirmation) Rolling percentiles must trail; fixed bands safe Hard label Default; VIX-band × slope × VRP set
HMM (filtered) ~1–3 days for persistent shifts; misses V-shapes Smoothed probabilities; refit on full history Posterior per state Cross-check on bands; P>0.7 gate
CUSUM / BOCPD 0–2 days for jumps; days for drifts Backtest alarms must use only past data Change event + probability Transition trigger / risk-off switch

In-Sample Fit vs Ex-Ante Usefulness

A high in-sample regime-classification accuracy is not evidence of tradability. Two failure modes recur:

  1. Smoothed/fitted labels: decoding the full series with the final fitted model gives every date a label that "knew" the future. Any backtest on those labels is lookahead.
  2. Surviving-threshold tuning: bands chosen because they historically bracket the data will be silently re-tuned after every crash.

The honest metric is: run the detector as it would have run daily (filtered outputs only, refits on trailing windows), then measure (a) detection delay per true regime change, (b) false-flip rate per year, (c) downstream P&L of the regime-gated rules vs the ungated baseline, net of costs per the EV contract.

Validation Approach

  • Point-in-time replay: regenerate all labels daily from data ≤ decision time; store labels, don't post-hoc classify.
  • Walk-forward: refit (HMM/CUSUM thresholds) on trailing windows only; freeze and log every refit for auditability, in the spirit of TOMIC's operational infrastructure6.
  • Regime-conditional baselines: per the EV contract, report VRP spread and a GARCH/HAR vol forecast 4 as the baseline within each regime; a regime rule earns its keep only if it beats the ungated baseline net of costs.
  • Stress the labels: perturb thresholds/parameters ±20% and check the downstream conclusion survives; if the edge exists only at one threshold value, it does not exist.
  • Delay accounting: for each historical regime change, record when each method actually flipped; publish the distribution, not the best case.

Links

References

  • Hamilton, J. (1989). "A New Approach to the Economic Analysis of Nonstationary Time Series and the Business Cycle." Econometrica 57(2), 357–384. https://doi.org/10.2307/2938279
  • Adams, R. & MacKay, D. (2007). "Bayesian Online Changepoint Detection." arXiv:0710.3742. https://arxiv.org/abs/0710.3742
  • Page, E. S. (1954). "Continuous Inspection Schemes." Biometrika 41(1–2), 100–115. https://doi.org/10.1093/biomet/41.1-2.100
  • Nystrup, P. et al. (2015). "Regime-based versus static asset allocation: Letting the data speak." Journal of Asset Management 16 — HMM filtering vs in-sample decoding pitfalls. https://doi.org/10.1057/jam.2015.19

  1. Hamilton's regime-switching framework; the filtered (not smoothed) probability is the only ex-ante quantity. ↩↩

  2. BOCPD gives online, causal run-length posteriors — detection delay is explicit and tunable. ↩

  3. CUSUM: latency vs false-alarm tradeoff set by the decision threshold via average run length. ↩

  4. EWMA/GARCH vol estimation as the forecast baseline in Hull VaR. ↩

  5. The books' qualitative condition language formalized by band rules in Natenberg volatility. ↩

  6. Auditability and operational discipline of a rules-based desk in TOMIC infrastructure. ↩