Summary¶
The system runs on two vendor feeds plus a Python pricing library: ThetaData for options data (historical 1-minute bars from 2016-01-01 and real-time OPRA, with roughly 2–4 concurrent request threads and no per-minute rate limit), Tiingo for underlying EOD OHLCV on its free tier (~1,000 requests/day), and py_vollib / py_vollib_vectorized for Black-Scholes-class pricing and greeks. These constraints directly shape backtest coverage, stress-test design, and feature-computation architecture. The books' infrastructure advice (redundancy, platform tooling) remains relevant for the execution layer 1.
Vendor Stack¶
| Component | Provider | Coverage | Key constraints |
|---|---|---|---|
| Options history | ThetaData Standard | 1-min option bars from 2016-01-01 | ~2–4 concurrent threads; no per-minute rate limit 2 |
| Options real-time | ThetaData (OPRA) | Streaming quotes/trades | Same thread budget shared with historical 2 |
| Options full depth | ThetaData Standard plan also covers underlyings | Underlying trades/quotes | One subscription serves both legs 2 |
| Underlying EOD | Tiingo free tier | Daily OHLCV, corporate-action adjusted | ~1,000 requests/day — batch EOD pulls, cache aggressively 3 |
| Pricing & greeks | py_vollib (+ py_vollib_vectorized) | BSM/European pricing, IV, vectorized greeks | CPU-only; American-style valuation needs care (binomial/Bjerksund or proxies) 4 |
Implications for Backtesting¶
- Backtests start 2016. Any option backtest window is bounded below by 2016-01-01; earlier claims require a different source or must be labeled as underlying-only approximations.
- Stress set = 2018, 2020, 2024. Vol-crisis stress testing uses: the Feb 2018 short-vol spike (Volmagedden), the COVID crash of Feb–Mar 2020, and the Aug 2024 VIX spike. These fall inside the data window and are the canonical regime stress episodes (see /60-regimes/index.md).
- Greeks come from the pricing library, not the vendor. Recompute greeks from py_vollib on stored bars so methodology is owned and reproducible in attested computations (/85-computations/index.md).
Implications for Feature Computation¶
- Thread limits (~2–4 concurrent) mean bulk historical pulls are latency-bound; design feature computation as batch jobs over cached local storage, not on-demand API fan-out.
- The no-per-minute-rate-limit property favors large sequential scans over parallel bursts; a simple queue with 2–4 workers saturates the allowance.
- Tiingo's ~1,000 req/day free tier rules out intraday underlying polling — restrict Tiingo to EOD universe maintenance; intraday underlying data comes via ThetaData.
- Precompute and persist greeks/skew/term-structure features per symbol-date so the EV pipeline never recomputes them at evaluation time.
Links¶
- EV Contract — every EV run consumes this stack
- Attested Computations
- Trading Plan
- TOMIC topic: Infrastructure
References¶
- ThetaData API documentation — plans, historical coverage from 2016, thread/rate limits: https://docs.thetadata.us/
- Tiingo pricing page — free tier limits and EOD API: https://www.tiingo.com/about/pricing
- py_vollib repository (Baruch College / vollib): https://github.com/vollib/py_vollib
- py_vollib_vectorized repository: https://github.com/vollib/py_vollib_vectorized