Frost Investment Research+886 2 2547 3268

Essay 02 · Frost Investment Research

Data leakage, catalogued: six routes by which the future contaminates a model

First published: October 2026 Editorial research — not investment advice

Leakage is any path by which information from outside the training sample — usually the future, often the label — enters model construction. Six named routes, each closable.

Leakage is the defining failure of applied machine learning in investment, and the reason is structural. A trading model predicts the future, so anything that lets a whisper of the future into training produces exaggerated statistics. Nothing afterwards rescues the model, because the flaw is in how information entered it, not in how cleverly it was fitted.

What follows is a field catalogue of the six routes seen most often in quantitative pipelines. Naming them matters: a bug you can name is a bug you can review for.

1. Target leakage

The feature is the label wearing a costume. A model of next-week returns that quietly includes a column computed from next week is the clown example, but real cases are subtler: a “rebalance flag” produced by the same simulated engine whose output is being predicted; a cost estimate whose formula embeds the fill that only existed post-trade. The tell is a feature too good to be true and too unrelated to be believed. The fix is to ask, for every column, exactly who could compute it and when.

2. Temporal leakage through random splits

Shuffle a time series and split it, and the training folds now contain Thursdays that have seen adjacent Fridays. Distance in time is information; randomisation destroys it. Later observations share volatility, regime, and calendar neighbourhood with earlier ones, and the model answers exam questions it sat next to. Correct practice keeps folds ordered in time — training strictly before the test slice, always.

3. Preprocessing leakage

The most common, the most innocent. Standardise, impute, or reduce dimensionality using the full history, then split — and the test slice has been touched by statistics that already knew it. Feature selection is the worst offender: rank predictors by their relationship with outcomes across the whole sample, and the selection step itself is a prophecy. Everything fitted must be fitted inside the training fold, then applied outward, never the reverse.

4. Leakage across fold boundaries (overlapping labels)

Financial labels often describe horizons longer than the sampling step: a five-day forward return beginning on day three spills across day eight. If day eight lands in the test fold, its outcome partially belongs to training. Industrial practice in financial machine learning purges training observations whose label windows overlap the test slice, and adds a short embargo after the split to stand still while autocorrelation drains. Skipping this produces suspiciously smooth out-of-sample curves.

5. Entity and repetition leakage

The same instrument appearing in both folds under different names — a ticker change, a dual listing, a spin-off — carries its own future across the boundary and out through the side door. Near-duplicate rows do the same quietly: resampled or overlapping windows that are, in substance, the same event counted twice. Duplicated information at the split boundary means the model has partly seen the test set before, however formal the split looks.

6. Peek-through metadata

Metadata is the softest leak. A column like “reporting gap” computed from when the next filing actually arrived encodes information a live system could not have at decision time. Membership flags, post-hoc industry classifications, and vendor fields carrying later revisions all smuggle hindsight through the side door of a tidy schema. The discipline is uniform: every feature must carry an honest timestamp — the moment a real system could first have read it.

The habit that closes all six

Each route is one violation of the same rule, and the rule fits in a sentence: at any simulated decision time t, use only the bytes a live process could have read at t. Data structures become point-in-time, joins become as-of, preprocessing becomes fold-local, labels become purged and embargoed, entities become deduplicated, metafields become honest.

Leakage is discovered, not avoided: the auditing mindset — suspicious columns, absurdly strong features, too-smooth curves — is the actual defence. The remainder is test protocol, which no model can grade itself on.

Frost Investment Research publishes free editorial essays about how machine learning gets evaluated in investing. Nothing here is personalised advice, and nothing on this site is for sale — including this line of work. Spotted an error, or want the next essay to cover something specific? Tell the editorial desk.

← Back to the article library