Synthetic data
Artificially generated data that preserves selected statistical properties of real data — used for augmenting sparse regimes, stress testing, and working around licensing or privacy constraints on the original.
The position
Synthetic market data is most attractive precisely where it is least trustworthy: the tails. Generators are fitted on history, so they reproduce the crises they were shown and not the one that arrives. Its honest use in systematic finance is adversarial — proving a strategy fails under regimes that never occurred — rather than generative. And there is an unresolved licensing question underneath: synthetic data derived from a licensed feed carries the original's terms far further than most agreements contemplate.
This section is a claim, not a definition. It is argued rather than asserted, and it is the part of this page you are invited to disagree with. Corrections and counter-arguments to [email protected] are published.