Essay 01 · Frost Investment Research
Why backtests flatter: five ways a simulation talks you into it
A backtest is an experiment on the one history we have. Five habits — from lookahead to frictionless fiction — bend that experiment toward the answer you were hoping for.
A backtest looks like science. It has data, a protocol, and a result printed in green. What it does not have — and cannot have — is a second history on which to replicate the result. There is exactly one past. An experiment that can never be repeated is fragile in a very specific way: nothing protects it from the small, quiet choices that bend it toward the answer you were already hoping for.
This essay is about five of those choices. None of them requires exotic mathematics. All of them survive review because each is boring, local, and individually defensible.
1. The future arrives early (look-ahead bias)
The classic form is a signal that uses information unavailable at the moment it is traded. A daily-close indicator is computed from the closing price, the simulator fills the order at that same closing price, and the trade is booked in the same line the signal was calculated from. The model has not predicted the close; it has read it.
Look-ahead also arrives through data hygiene. Restated earnings rest in the database with today’s values but yesterday’s dates — as if a trading system in 2019 could somehow have owned the corrected figures. Index membership lists ship “as of now”, quietly handing the model the future composition of the portfolio. Each is a small courtesy to the simulated trader that no live trader receives.
2. The sample forgets the dead (survivorship bias)
Test a strategy on today’s index constituents and you have removed every company that failed, merged, or was removed from the index — precisely the cases where things go wrong. The equity curve now rides a sample that was selected, retroactively, for success.
Survivorship is not an exotic trap; it is the default behaviour of convenient datasets. List files are easy, point-in-time histories are expensive, and the difference between them compounding silently inside every return figure. A related version hides in fund databases: only the funds that still exist report performance. Gaps of that kind are invisible on a chart and fatal in production.
3. Enough attempts, and luck looks like skill (multiple testing)
Search forty parameter combinations against the same decade and pick the best, and you have not found an edge — you have found the luckiest of forty. The past, unlike the future, is finite and fixed: it has a limited number of accidents, and a patient optimiser can tune a curve to echo them exactly.
The scientific failure is not the search itself. Searching is normal engineering. The failure is reporting the winner alone, without telling anyone a tournament took place. A deflation of this kind — showing the distribution of losers next to the chosen survivor — is the honest shape of that work.
4. The frictionless fiction
Simulated fills at the close, at the ask, at any size, always borrowable, never throttled by liquidity. Each assumption is a line of politeness written by the simulator. Aggregate them and an ordinary cost profile can evaporate the whole of a thin backtested edge. The uncomfortable result — that after costs the strategy resembles the index — is the one the simulation was never forced to confront.
5. One decade is not the world (regime narrowness)
A model tested on 2010–2025 has met exactly one rate environment, one liquidity regime, and one set of correlations. It is an expert on a decade, not on markets. Backtests are silent on the regimes they never saw, and the confusion of “worked recently, here” with “works” is the most common translation error in this business.
What an honest evaluation keeps
None of this makes simulation useless; it makes simulation provisional. The practices that keep it honest are consistent and few: keep processing inside the fold it belongs to; embargo the windows just after training so labels cannot peek; report every attempt, not the winner; charge realistic costs; prefer parameter plateaus over sharp peaks; and treat the final curve as a hypothesis that survived questioning — never as a promise.
A backtest, read correctly, is not a forecast. It is the first hostile question you ask of an idea before the market asks harder ones.
Frost Investment Research publishes free editorial essays about how machine learning gets evaluated in investing. Nothing here is personalised advice, and nothing on this site is for sale — including this line of work. Spotted an error, or want the next essay to cover something specific? Tell the editorial desk.
← Back to the article library