Spurious Predictability — Falsification Audit and Inflation Ledger
A walk-forward winner can be large because the pipeline invents signal, because the search picked a lucky specification, or because both happened. Sotirios D. Nikolopoulos, Spurious Predictability in Financial Machine Learning (arXiv:2604.15531), splits those failures into an order of operations. First falsify the whole workflow on induced nulls and on a microstructure placebo. Then, only if it survives, write down an inflation ledger: the gap between the optimized in-sample winner statistic and the same configuration’s disjoint walk-forward statistic, with effective multiplicity beside it. This page is that protocol as a guide. It sits next to AlgoXpert’s IS–WFA–OOS gates, nested CV, and Deflated Sharpe. It is not a deployment recipe and not a signal.
The paper’s own rank matches the house stance. Clearing the audit is a necessary condition for taking a backtest seriously. It is not evidence of economic edge, not a cost model, and not a certificate. Section 4.5 says so in those terms. A shippable validated edge remains rare; this audit is one way a candidate fails in public before anyone argues about its Sharpe.
1 / What the object actually is
Global null. Excess returns are a martingale difference: E[rₜ | Fₜ₋₁] = 0. A fixed signal measurable at t−1 then has conditional mean zero P&L. The paper’s claim is that adaptive search and a leaky workflow can still print a significant statistic under that null.
Two phases, required to be disjoint. In-sample selection versus walk-forward evaluation, intersection empty. Disjoint means the entire map — preprocessing, features, tuning, model choice — is measurable on the in-sample σ-field only. Studentized means use a Newey–West HAC variance with Bartlett kernel and the rule-of-thumb bandwidth Lₙ = ⌊4(n/100)^(2/9)⌋. The paper warns that for a short walk-forward (it gives n_WF = 252 days as the example) those t-statistics have heavy tails and over-reject. Used as an admission gate, that bias makes the screen harsher on acceptance, not kinder.
Effective multiplicity. For the correlation matrix Σ of the in-sample statistics,
K_eff = (tr Σ)² / ‖Σ‖²_F = K² / ‖Σ‖²_F.
Independence gives K_eff = K. Collinearity pulls K_eff toward 1. High dimension: linear shrinkage toward the identity (Schäfer–Strimmer). Proposition 1, under a structured high-correlation geometry and K_eff → ∞: the expected absolute in-sample winner is Θ(√log K_eff). A disjoint walk-forward evaluation of the already chosen index stays O_p(1) under the null (Lemma 1). The primary ledger entry is the absolute gap
ΔZ ≡ Z*ᵢₛ − Z*_WF,
with Z*ᵢₛ = max |Z_IS,k| and Z*_WF = |Z_WF| of that same selected k. The ratio BIF_raw = Z*ᵢₛ / Z*_WF is undefined at zero. The stabilized version floors the denominator at τ = 0.5. The paper is explicit that BIF^stab_τ is an intra-study ranking heuristic, not a cross-study number, because it depends on τ. Cross-study comparison is supposed to use ΔZ.
Stage 1 — falsification. Run the unchanged pipeline through Monte Carlo draws of induced nulls. For each environment e, compare the observed walk-forward winner to the empirical (1 − α/5) quantile of that environment’s null winners (Bonferroni across five environments). Falsify if any environment exceeds its quantile. The paper says this controls family-wise error and spends power; a fail can be a finite-sample fluke rather than a leak.
Five environments in Section 4.2:
| Env | What it is | What a “pass” is refusing |
|---|---|---|
| A | i.i.d. Gaussian noise, zero mean | Pure selection painted as skill |
| B | Two-state Markov volatility, zero conditional mean | Scale or normalization rules that trade a vol state |
| C | Not an MDS in observed returns. Latent efficient return is an MDS; observed return adds Roll-style bid–ask bounce (MA(1), negative autocorrelation) | Monetizing microstructure by a bad lag or index |
| D | One mean-zero factor plus idiosyncratic noise | Reading factor exposure as alpha |
| E | GARCH(1,1), zero conditional mean | HAC and risk adjustment under vol clustering |
Appendix B defaults that were actually in the fetched text: Environment A annualized vol 0.20 so daily σ = 0.20/√252. Environment B low-regime annualized vol 0.10, high-regime multiplier 3.0 (so 0.30), persistence p₁₁ = p₂₂ = 0.98. The fetch’s HTML conversion cut off inside Appendix B.3 (Environment C’s full algebra). Body text still states the construction: latent MDS plus Roll MA(1). Section 4.6’s blind-parameter example draws the MA coefficient from Uniform[−0.8, −0.2] instead of fixing θ = −0.5, with a public Θ_dev and a withheld Θ_audit, because a static placebo can be overfit too.
Stage 2 — inflation class on the target data. Only after Stage 1. Same workflow, strict disjoint split, report Z*ᵢₛ, Z*_WF, K̂_eff, ΔZ, and BIF^stab as a side heuristic. Flag “null-implausible inflation” if ΔZ exceeds the 0.99 quantile of a pipeline-specific null for ΔZ. Appendix thresholds are called illustrative. If you cannot simulate your own (K_eff, n_IS, n_WF), the paper says use a conservative worst case at maximal effective multiplicity. That is an instruction to not copy a printed cutoff from someone else’s grid.
The reference implementation named in the paper is an R package, QuantAudit. As of the text we read, it is “prepared for public release” upon journal publication. This draft does not assume you can install it today, and we did not run it.
2 / What it does not claim — and where it fails
- Pass ≠ edge. Section 4.5: robustness to a specified set of artifacts. The book can still be economically small, concentrated, or dead after costs. The audit is a stress test.
- Power collapses on rare signals. The positive control (Section 5.6) is a threshold autoregression rₜ = φ rₜ₋₁ 1{|rₜ₋₁| > θσ} + ε, annualized vol fixed at 15%, T = 2,520, 60/40 split, N = 1,000 per cell. The workflow picks a breakout threshold from {0.5, 0.7, …, 3.5} by absolute in-sample t, at least 10 in-sample trades, and “detects” if |Z_WF| > 1.96. At φ = 0.05 and θ = 3.0 the printed detection rate is 4.8%, at or under the nominal 5% size. At φ = 0.25 and θ = 1.0 it is 100%. The audit will miss a real mechanism that almost never turns on inside the walk-forward window. That is a designed bias toward Type I control, and it will false-kill sparse trend rules. Leave that possibility open; do not treat a Stage 1 fail on a rare-breakout book as proof the book was leakage.
- HAC over-rejection means Stage 1 can call noise “significant OOS.” The paper prefers the induced-null quantiles over the N(0,1) critical value for that reason. A reader who thresholds |Z| > 1.96 on 252 HAC days is using the approximation the authors distrust.
- K_eff is not a universal extreme-value parameter. Proposition 1 is tied to a packing geometry. Table 4 shows the EVT approximation breaking when one noisy cluster dominates: at ρ = 0.60, 1 cluster, K = 500, K̂_eff = 2.8, observed E[|Z*ᵢₛ|] = 2.56 against an EVT prediction of 1.34 (error 91.5%). Low K_eff does not mean “no inflation,” which Section 5.3 then shows on purpose: XGBoost hyperparameter tuning at K = 400 under white noise has K̂_eff = 2.93 but mean Z*ᵢₛ = 2.56, ΔZ = 1.77, in-sample fail rate 88.2%, while walk-forward fail rate stays 4.7%. Correlation of raw scores is not correlation of thresholded positions (Table 6: at K = 1,000 and ρ = 0.9, K_eff of predictions stays 1.24; after a fixed threshold of 2.0, signal K_eff rises to 3.34, amplification 2.68).
- BIF numbers inside the paper do not all agree. Table 1, Environment A, T = 2,520, N = 1,000, lists median BIF^stab_0.5 (on the subset |Z*_WF| ≥ 0.5) at K = 1,000 as 3.33. The next paragraph says the median BIF at K = 1,000 on the reported subset is 5.12. Both sentences were on the page we read. Do not average them into a cleaner factor. A teardown uses ΔZ, which the same table prints as 2.68 at that row, and treats the ratio as the unstable object the authors already demoted.
- Empirical cases are demonstrations, not a prevalence. Section 6 says they are not an estimate of how often published ML is spurious.
- Meta-overfitting. A pipeline can be tuned to pass a fixed placebo. The blind-parameter split is the paper’s patch, not something a single public notebook automatically has.
- We did not re-simulate. Every Monte Carlo figure below is the paper’s. If a later clean room disagrees, the clean room is the check; this page is not that check.
3 / The ledger, only where we read the cells
Table 1, white-noise selection, T = 2,520, means unless noted. Winner FWP is the family-wise false-positive rate of the selected winner at |Z| > 1.96. BIF column is the median on |Z*_WF| ≥ 0.5.
| K | Mean Z*ᵢₛ | Mean Z*_WF | ΔZ | BIF^stab_0.5 | Winner FWP % |
|---|---|---|---|---|---|
| 1 | 0.81 | 0.77 | 0.04 | 0.65 | 5.3 |
| 10 | 1.90 | 0.81 | 1.09 | 1.80 | 40.2 |
| 100 | 2.79 | 0.82 | 1.97 | 2.60 | 99.9 |
| 1000 | 3.48 | 0.80 | 2.68 | 3.33 | 100.0 |
Walk-forward means stay near 0.8, which the text compares to the half-normal expectation √(2/π) ≈ 0.80. In-sample means climb. That is the picture. The K = 5, 50, 200, 500 rows exist in the paper and are omitted here only for length; they are monotone in the same direction. The 5.12 sentence is the conflict noted above, not an extra row.
Table 5, white noise, K = 400, boosting, N = 1,000 (means; BIF median on the subset):
| Workflow | K̂_eff | Mean Z*ᵢₛ | Mean Z*_WF | ΔZ | IS fail % | WF fail % |
|---|---|---|---|---|---|---|
| FeatureMining | 400.00 | 3.21 | 0.84 | 2.37 | 100.0 | 6.6 |
| Hyperparameter tuning | 2.93 | 2.56 | 0.79 | 1.77 | 88.2 | 4.7 |
| TrendFollowing | 5.86 | 2.19 | 0.78 | 1.41 | 64.2 | 5.0 |
Elastic-net tuning, same K = 400 white-noise control (Table 8): K̂_eff 1.727, mean |Z*ᵢₛ| 1.817, mean |Z*_WF| 0.808, mean ΔZ 1.009, IS fail 35.8%, WF fail 5.4%. Hundreds of λ values compressed; inflation shrank and did not disappear.
Two market demonstrations, gross of costs where the paper says so:
- SPY deep-learning direction, strict temporal split, internal pre-2016 validation, no class rebalancing. Post-2015 annualized alpha versus buy-and-hold 0.39%, p = 0.931. Trigger rate 69.4% in calm vs 88.7% in high vol; boundary-mass statistic 0.174 vs 0.066 (ε = 0.01; regime cut = training 80th percentile of lagged vol). On N = 30 zero-mean GARCH(1,1) paths the same pipeline’s mean trigger rates were 68.9% calm and 79.1% high vol; mean annualized alpha −0.17%; mean gross Z −0.014. The paper’s reading: a volatility-responsive activation rule, not a directional model. Gross of costs, so the economic failure is understated.
- FF25 ridge, λ = 0.01, expanding window, long top three and short bottom three, 1991–2025, 419 OOS months, Fama–French missing code −99.99 scrubbed. Gross HAC Z 3.153 (p = 0.002); FF3 alpha Z 2.708 (p = 0.007); BIF defined there as |Z_gross| / max(|Z_α|, τ) = 1.16. In this sample the beta-confound diagnostic did not fire. The paper says that outright. A guide that claimed “ML on FF25 is only beta” would be inventing a result the case study refused.
- SPY log realized variance, HAR vs Lasso. HAR OOS R² 0.238 (RMSE 0.000584). Lasso + random CV 0.212 (RMSE 0.000446). Lasso + walk-forward refit every 5 days 0.211 (RMSE 0.000594). VIR = R²_random / R²_WF = 1.006. ΔR² = 0.001. The vol-forecast case is a non-inflation result. The methodological lesson in the paper is that the size of validation inflation has to be measured per pipeline, not assumed.
4 / How a reader could tear it down next
Not an order. A next attempt on one frozen workflow:
- Write the workflow down before any null seed: features, scaler (expanding or full-sample — full-sample is the leak the taxonomy names), label horizon, execution lag, search space K, selection metric, IS/WF cut. If the scaler is full-sample, Stage 1 should be able to falsify you. If it cannot, your null is not stressing the leak you think it is.
- Stage 1 on A–E with your own draws, not with a copied Z cutoff. Include the bounce placebo (Environment C) if the book trades close-to-close on bid/ask contaminated prices. Keep a parameter slice you did not tune on (the paper’s Θ_audit idea) or you will overfit the audit.
- Stage 2 ledger is ΔZ and K_eff of the thing you trade, meaning the position after the threshold, not the raw score. Table 6 is the reason. Report BIF^stab_0.5 only inside this study, and do not compare it with a paper that used another τ.
- Do not stop at a pass. A survivor still needs costs, a causal outer segment (Li’s trinity inequality, this slate’s Learn draft; nested CV), and a multiplicity adjustment aimed at a performance statistic (DSR). This audit’s Z is a HAC t on mean return. It is not a Sharpe deflator. They answer different questions; run the one your claim requires.
- If you fail only in the sparse corner of a positive control you actually have, record the activation rate. The paper’s 4.8% cell is a warning that “falsified” can mean “the walk-forward window almost never saw the state,” which is a different bug from lookahead.
5 / Verdict
The protocol is a gate plus a ledger, in that order. Induced nulls (zero mean, and one bounce placebo that is not zero mean in the observed price) ask whether the pipeline can hallucinate a walk-forward winner. ΔZ asks how far the in-sample max ran ahead of a disjoint evaluation, with K_eff so that a correlated grid is not billed as K independent bets. BIF is the unstable ratio the paper tells you not to carry across studies. None of this is a strategy, and a pass is not a ship. The FF25 case keeping its alpha, and the volatility case showing essentially no random-CV inflation, are there so the audit is not a machine that always says no. The white-noise tables are there so a K = 100 “significant” in-sample winner is not a surprise. Next attempt: your pipeline, your seeds, ΔZ on positions, then costs and DSR — still untested on this site.
Source
- Sotirios D. Nikolopoulos, Spurious Predictability in Financial Machine Learning. https://arxiv.org/abs/2604.15531
- Read 2026-10-03 from the arXiv HTML full text through Appendix A, Appendix B.1–B.2, and the start of B.3. Appendix B.3 onward (full Environment C/D/E algebra and later robustness appendices) was truncated by the fetch. Environment definitions used above are from Section 4 and the appendix defaults that rendered. JEL on the paper: C52, C53, C55, G17.
Numbers
Every statistic in the tables and case studies above is printed in the paper and attributed to the paper’s Monte Carlo or to the paper’s market demonstration. In particular: τ = 0.5; Bonferroni 1 − α/5 across five environments; ΔZ flag at the 0.99 quantile; Table 1 T = 2520, N = 1000; the 5.12 vs 3.33 BIF conflict at K = 1000; Table 4 error 91.5% at the one-cluster boundary; Table 5 and Table 8 at K = 400; detection 4.8% and 100% at the cited (φ, θ) corners; SPY alpha 0.39% (p 0.931), GARCH-null N = 30; FF25 Z 3.153 / alpha Z 2.708 / BIF 1.16 / 419 months; VIR 1.006. No Quant for Free rerun. summary.json not yet — no clean-room run.
summary.json. Figures in the text are the source’s, labeled as such. See the full disclaimer.