Skip to content

Summary

The system runs on two vendor feeds plus a Python pricing library: ThetaData for options data (historical 1-minute bars from 2016-01-01 and real-time OPRA, with roughly 2–4 concurrent request threads and no per-minute rate limit), Tiingo for underlying EOD OHLCV on its free tier (~1,000 requests/day), and py_vollib / py_vollib_vectorized for Black-Scholes-class pricing and greeks. These constraints directly shape backtest coverage, stress-test design, and feature-computation architecture. The books' infrastructure advice (redundancy, platform tooling) remains relevant for the execution layer 1.

Vendor Stack

Component Provider Coverage Key constraints
Options history ThetaData Standard 1-min option bars from 2016-01-01 ~2–4 concurrent threads; no per-minute rate limit 2
Options real-time ThetaData (OPRA) Streaming quotes/trades Same thread budget shared with historical 2
Options full depth ThetaData Standard plan also covers underlyings Underlying trades/quotes One subscription serves both legs 2
Underlying EOD Tiingo free tier Daily OHLCV, corporate-action adjusted ~1,000 requests/day — batch EOD pulls, cache aggressively 3
Pricing & greeks py_vollib (+ py_vollib_vectorized) BSM/European pricing, IV, vectorized greeks CPU-only; American-style valuation needs care (binomial/Bjerksund or proxies) 4

Implications for Backtesting

  • Backtests start 2016. Any option backtest window is bounded below by 2016-01-01; earlier claims require a different source or must be labeled as underlying-only approximations.
  • Stress set = 2018, 2020, 2024. Vol-crisis stress testing uses: the Feb 2018 short-vol spike (Volmagedden), the COVID crash of Feb–Mar 2020, and the Aug 2024 VIX spike. These fall inside the data window and are the canonical regime stress episodes (see /60-regimes/index.md).
  • Greeks come from the pricing library, not the vendor. Recompute greeks from py_vollib on stored bars so methodology is owned and reproducible in attested computations (/85-computations/index.md).

Implications for Feature Computation

  • Thread limits (~2–4 concurrent) mean bulk historical pulls are latency-bound; design feature computation as batch jobs over cached local storage, not on-demand API fan-out.
  • The no-per-minute-rate-limit property favors large sequential scans over parallel bursts; a simple queue with 2–4 workers saturates the allowance.
  • Tiingo's ~1,000 req/day free tier rules out intraday underlying polling — restrict Tiingo to EOD universe maintenance; intraday underlying data comes via ThetaData.
  • Precompute and persist greeks/skew/term-structure features per symbol-date so the EV pipeline never recomputes them at evaluation time.

Links

References

  • ThetaData API documentation — plans, historical coverage from 2016, thread/rate limits: https://docs.thetadata.us/
  • Tiingo pricing page — free tier limits and EOD API: https://www.tiingo.com/about/pricing
  • py_vollib repository (Baruch College / vollib): https://github.com/vollib/py_vollib
  • py_vollib_vectorized repository: https://github.com/vollib/py_vollib_vectorized

Source Notes


  1. The Option Trader's Hedge Fund, infrastructure topic (sources[]: tomic-infra). ↩

  2. ThetaData documentation (sources[]: thetadata-docs). ↩↩↩

  3. Tiingo pricing docs (sources[]: tiingo-docs). ↩

  4. py_vollib and py_vollib_vectorized repositories (sources[]: pyvollib, pyvollib-vectorized). ↩