The Impossible Trinity of Time-Series Validation
Walk-forward, nested CV, and the Deflated Sharpe Ratio answer different complaints about the same backtest. Jiayu Li’s September 2026 note, The Impossible Trinity of Time-Series Validation (arXiv:2609.29530), claims those complaints are not three tastes. They are one budget. You can ask a split to train on most of the sample, to test most of the sample, and to train only on the past. The paper says you cannot have all three, and it writes the bill as an inequality. This page is the companion theory next to Walk-Forward Optimization, Nested CV, Deflated Sharpe, and look-ahead bias. It is not another calibrator and not a new Sharpe deflator.
The inequality is in the paper. It is not a slogan we added.
1 / What the object actually is
A validation scheme is a list of folds. Fold i trains on a set Rᵢ and tests on a nonempty Eᵢ, and the two sets do not overlap. Three coordinates:
- α — worst-fold training fraction. The smallest |Rᵢ| / T. Sufficiency: the evaluated model should look like the one you would train on all T.
- β — test coverage. The fraction of the sample that appears in at least one test set.
- Causality, in two numbers. For a test point t, Fᵢ(t) is the training data that sits after t. Λ is the anti-causal mass: the largest such future-training block, as a fraction of T. δ is the anti-causal margin: how far you must walk forward from a test point before you hit a training point. Strict causality is Λ = 0 and δ = ∞.
Theorem 1 then says, with no statistical assumption on parts (a)–(c), that every scheme on a line of length T satisfies
α + β ≤ 1 + Λ
and
α + min{β, δ/T} ≤ 1.
Going past the causal frontier α + β = 1 means you trained on the future. That future data cannot be parked far away: if Λ > 0 then some test point has training within (1 − α)T steps after it. The harm of that leak, under β-mixing, is bounded by distance, not by volume: at a test point the leakage bias is at most 2M β_mix(δ), where losses are bounded by M. Volume of future data is the ledger entry. Proximity is what the mixing coefficient prices.
Corollary 3 is the sentence this site’s walk-forward guide has been circling: expanding walk-forward is exactly the Pareto frontier of strictly causal schemes. Any (α, β) with α + β ≤ 1 is weakly dominated by an expanding walk-forward. Last-block hold-out is the one-block member of that same family. At equal coverage they share the same worst-fold α. What separates them is average training size. With m equal blocks, the mean training fraction of expanding walk-forward approaches 1 − β/2, while a single hold-out sits at 1 − β. Theorem 4 adds the sting: if the test blocks tile the whole sample (β = 1), even the average training fraction of a strictly causal scheme is at most (m − 1)/(2m), which is less than one half.
Purged k-fold does not break the ledger. Contiguous k-fold has α = (k − 1)/k, β = 1, Λ = (k − 1)/k, and δ = 1 at the block edge — it buys the minimum Λ its coordinates require, all of it at the closest possible margin. An embargo spends sufficiency to push δ out to h. If the process forgets quickly, β_mix(h) is cheap. If the process is non-stationary, the embargo does not buy the part of causality that matters: a model that trained on the future of a test point has seen the regime the deployed model will not have seen. The paper’s line is that only walk-forward-type schemes pay that bill in full.
Label horizon H is the same debit. Overlapping labels force a gap. The combinatorial statements are written for H = 0; the correction is O(mH/T), taken out of sufficiency. That is the same purging motive as López de Prado (2018), priced here as a term in the inequality rather than as a folklore warning.
Section 5.2 then points at the tools this site already ships. Tuning takes a max over N configurations. Low coverage inflates the winner (variance of order 1/(βT) times a √(2 ln N) selection term, in the paper’s sketch). Leakage bias is per configuration, and the max prefers the configuration that exploited the leak. Learning curves of different hyperparameters have different slopes, so a ranking at αT need not be the ranking at T. The prescribed shape is the one nested CV already implements: inner selection, outer strictly causal segment never touched by selection, then a multiplicity correction — Deflated Sharpe or PBO — on the winner.
2 / What it does not claim — and where it fails
This is a conservation law for splits, plus a price list. It does not say a strategy has edge, and it does not say walk-forward is the most powerful test.
- The trinity disappears in the i.i.d. limit. If the augmented sample is i.i.d. and H = 0, β_mix(d) = 0 for every d ≥ 1. Shuffled k-fold is then legitimate. The paper is explicit: the trinity is created by dependence and non-stationarity, not by the words “time series.” Bergmeir, Hyndman, and Koo (2018) on autoregressions with uncorrelated errors is treated as a small algorithm-specific exchange rate, not as a counterexample.
- Pessimism is not neutrality. Under a monotone learning curve, a causal scheme is biased toward higher loss (Proposition 6). That bias favors simpler models when you compare at a small training size. “Walk-forward is conservative” can still pick the wrong horse.
- Coverage has a regime term, not just a variance term. Proposition 7(ii): an untested interval can hide a risk gap Δ, and no estimator built only on the covered loss record can see it. Worst-case error at least (1 − β)Δ/2. CPCV is discussed as more paths through regimes, paid for with less sufficiency and more computation — not as an escape from the ledger.
- The numerical illustration is pure noise, on purpose. It is not a market study. Do not quote its information coefficients as evidence about BTC, ES, or any book.
- Mixing bounds need a mixing process and bounded loss. Long memory and discrete-innovation autoregressions that do not mix are acknowledged; Theorem 5 moves the exchange rate to a weaker metric, but that rate is not something you read off a Sharpe printout.
- Non-monotone learning curves (double descent) weaken Proposition 6 to a min/max of L on [αT, T]. The paper says so. A slogan that “less training always looks worse” is stronger than the theorem.
- Practical card item 1 names c ≈ 2–3 multiples of the dependence scale for gaps. That is guidance in Section 6, not a theorem.
A scheme can sit on the causal frontier and still be a p-hacked window. The trinity does not replace Deflated Sharpe. It tells you which coordinate you spent before you deflate.
3 / How a reader could tear it down next
Not a trading rule. A next attempt on a scheme you already run:
- Write the five coordinates for the scheme you actually ship: α, mean α, β, Λ, δ. One line. If you cannot name δ relative to label horizon H and a dependence scale, you do not know whether your Λ > 0 is harmless.
- Check the frontier arithmetic. If the scheme is strictly causal, confirm α + β ≤ 1 after the H-gap is removed from training. If you are above 1, the excess is a lower bound on Λ. That is Theorem 1(a), and it does not require a p-value.
- Separate volume from severity on one null. The paper’s Section 5.3 design is reproducible as a method check, not as a market claim: i.i.d. noise, overlapping labels of horizon H, a learner that can latch onto a near neighbour, shuffled k-fold versus contiguous k-fold versus a purged gap versus expanding walk-forward. Same (α, β, Λ) for shuffled and contiguous; the reported association should collapse as the margin grows. If your pipeline’s “OOS IC” does not collapse on that null, the leak is in the pipeline.
- Then put the outer segment where Corollary 3 puts it. Inner loop may be purged k-fold if you believe the memory is short and the sanitization overhead k(2H+h)/T is small. The number you would quote as live performance comes from a strictly causal outer segment that never entered selection. Deflate N the way the DSR page already describes.
- Do not “solve” the trinity by shuffling. Section 6 calls shuffled k-fold unusable once labels overlap or residuals are dependent. The illustration below is why.
4 / The illustration the paper actually prints
Section 5.3, labeled as a distinction between Λ and effective leakage, not as a market result. Design, as printed: T = 3000; returns i.i.d. standard normal; labels are H = 20 period forward sums (true predictability zero); four EMAs of returns (half-lives 5/10/20/40); 1-nearest-neighbour regression; metric is the correlation of prediction and label (information coefficient). Eight seeds.
Exclusion radius g (training forbidden inside distance g of the test point). Reported IC:
| g | 0 | 2 | 4 | 6 | 8 | 10 | 12 | 14 | ≥ 16 |
|---|---|---|---|---|---|---|---|---|---|
| IC | +0.340 | +0.134 | +0.071 | +0.040 | +0.023 | +0.012 | +0.004 | −0.000 | ≈ 0 |
Named schemes on the same noise (true value 0), reported IC ± s.e. as printed:
| Scheme | Reported IC |
|---|---|
| Shuffled 5-fold | +0.318 ± 0.012 |
| Contiguous 5-fold | +0.004 ± 0.012 |
| Purged 5-fold (gap = H) | +0.000 ± 0.012 |
| Walk-forward (gap = H) | +0.017 ± 0.015 |
| Last-20% hold-out (gap = H) | +0.044 ± 0.028 |
The abstract’s round numbers match that table’s first two rows: shuffled 5-fold +0.32, contiguous 5-fold +0.004, “using the same amount of future data.” The walk-forward in the table is expanding, last half of the sample as ten equal test blocks (initial window w = 0.5), with an H-gap. Hold-out standard error is the largest, which is the variance price of small β in this draw — a property of this illustration, not a universal ranking of hold-out versus walk-forward.
5 / Verdict
The inequality is real on the page: α + β ≤ 1 + Λ, and with the margin, α + min{β, δ/T} ≤ 1. Expanding walk-forward is the causal Pareto frontier, not a house style. k-fold buys the future; an embargo buys distance and spends sufficiency; non-stationarity is not redeemed by distance. That does not make every walk-forward honest, and it does not make every purged CV a leak. It makes the unpaid coordinate visible.
A shippable validated edge is still rare. This paper does not supply one. It supplies the ledger you attach to the next attempt: print (α, α-bar, β, Λ, δ), put any quoted live number on a strictly causal outer segment, and deflate the search. If those coordinates are missing, the backtest has not failed yet. It has not been specified.
Source
- Jiayu Li, The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality, September 2026. https://arxiv.org/abs/2609.29530
- Read 2026-10-03 from the arXiv HTML full text (abstract through Appendix A and the reference list). Theorems, tables, and the Section 5.3 illustration above are from that text.
Numbers
All figures above are the paper’s, attributed in place: the two inequalities; k-fold coordinates (k−1)/k; the half-sample bound on mean training under full causal coverage; T = 3000, H = 20, 8 seeds; the g-sweep ICs; the five named-scheme ICs with standard errors; Section 6’s gap guidance c ≈ 2–3 (guidance, not a theorem); the selection sketch ς√(2 ln N). No Quant for Free backtest. No market Sharpe. summary.json not yet — no clean-room run — this is a theory note, not a strategy teardown.
summary.json. Figures in the text are the source’s, labeled as such. See the full disclaimer.