Quant for Free
Home / Guides / Spurious Predictability
Guide · falsification audit · not a deployment recipe

Spurious Predictability — Falsification Audit and Inflation Ledger

A walk-forward winner can be large because the pipeline invents signal, because the search picked a lucky specification, or because both happened. Sotirios D. Nikolopoulos, Spurious Predictability in Financial Machine Learning (arXiv:2604.15531), splits those failures into an order of operations. First falsify the whole workflow on induced nulls and on a microstructure placebo. Then, only if it survives, write down an inflation ledger: the gap between the optimized in-sample winner statistic and the same configuration’s disjoint walk-forward statistic, with effective multiplicity beside it. This page is that protocol as a guide. It sits next to AlgoXpert’s IS–WFA–OOS gates, nested CV, and Deflated Sharpe. It is not a deployment recipe and not a signal.

The paper’s own rank matches the house stance. Clearing the audit is a necessary condition for taking a backtest seriously. It is not evidence of economic edge, not a cost model, and not a certificate. Section 4.5 says so in those terms. A shippable validated edge remains rare; this audit is one way a candidate fails in public before anyone argues about its Sharpe.

1 / What the object actually is

Global null. Excess returns are a martingale difference: E[rₜ | Fₜ₋₁] = 0. A fixed signal measurable at t−1 then has conditional mean zero P&L. The paper’s claim is that adaptive search and a leaky workflow can still print a significant statistic under that null.

Two phases, required to be disjoint. In-sample selection versus walk-forward evaluation, intersection empty. Disjoint means the entire map — preprocessing, features, tuning, model choice — is measurable on the in-sample σ-field only. Studentized means use a Newey–West HAC variance with Bartlett kernel and the rule-of-thumb bandwidth Lₙ = ⌊4(n/100)^(2/9)⌋. The paper warns that for a short walk-forward (it gives n_WF = 252 days as the example) those t-statistics have heavy tails and over-reject. Used as an admission gate, that bias makes the screen harsher on acceptance, not kinder.

Effective multiplicity. For the correlation matrix Σ of the in-sample statistics,

K_eff = (tr Σ)² / ‖Σ‖²_F = K² / ‖Σ‖²_F.

Independence gives K_eff = K. Collinearity pulls K_eff toward 1. High dimension: linear shrinkage toward the identity (Schäfer–Strimmer). Proposition 1, under a structured high-correlation geometry and K_eff → ∞: the expected absolute in-sample winner is Θ(√log K_eff). A disjoint walk-forward evaluation of the already chosen index stays O_p(1) under the null (Lemma 1). The primary ledger entry is the absolute gap

ΔZ ≡ Z*ᵢₛ − Z*_WF,

with Z*ᵢₛ = max |Z_IS,k| and Z*_WF = |Z_WF| of that same selected k. The ratio BIF_raw = Z*ᵢₛ / Z*_WF is undefined at zero. The stabilized version floors the denominator at τ = 0.5. The paper is explicit that BIF^stab_τ is an intra-study ranking heuristic, not a cross-study number, because it depends on τ. Cross-study comparison is supposed to use ΔZ.

Stage 1 — falsification. Run the unchanged pipeline through Monte Carlo draws of induced nulls. For each environment e, compare the observed walk-forward winner to the empirical (1 − α/5) quantile of that environment’s null winners (Bonferroni across five environments). Falsify if any environment exceeds its quantile. The paper says this controls family-wise error and spends power; a fail can be a finite-sample fluke rather than a leak.

Five environments in Section 4.2:

Env What it is What a “pass” is refusing
A i.i.d. Gaussian noise, zero mean Pure selection painted as skill
B Two-state Markov volatility, zero conditional mean Scale or normalization rules that trade a vol state
C Not an MDS in observed returns. Latent efficient return is an MDS; observed return adds Roll-style bid–ask bounce (MA(1), negative autocorrelation) Monetizing microstructure by a bad lag or index
D One mean-zero factor plus idiosyncratic noise Reading factor exposure as alpha
E GARCH(1,1), zero conditional mean HAC and risk adjustment under vol clustering

Appendix B defaults that were actually in the fetched text: Environment A annualized vol 0.20 so daily σ = 0.20/√252. Environment B low-regime annualized vol 0.10, high-regime multiplier 3.0 (so 0.30), persistence p₁₁ = p₂₂ = 0.98. The fetch’s HTML conversion cut off inside Appendix B.3 (Environment C’s full algebra). Body text still states the construction: latent MDS plus Roll MA(1). Section 4.6’s blind-parameter example draws the MA coefficient from Uniform[−0.8, −0.2] instead of fixing θ = −0.5, with a public Θ_dev and a withheld Θ_audit, because a static placebo can be overfit too.

Stage 2 — inflation class on the target data. Only after Stage 1. Same workflow, strict disjoint split, report Z*ᵢₛ, Z*_WF, K̂_eff, ΔZ, and BIF^stab as a side heuristic. Flag “null-implausible inflation” if ΔZ exceeds the 0.99 quantile of a pipeline-specific null for ΔZ. Appendix thresholds are called illustrative. If you cannot simulate your own (K_eff, n_IS, n_WF), the paper says use a conservative worst case at maximal effective multiplicity. That is an instruction to not copy a printed cutoff from someone else’s grid.

The reference implementation named in the paper is an R package, QuantAudit. As of the text we read, it is “prepared for public release” upon journal publication. This draft does not assume you can install it today, and we did not run it.

2 / What it does not claim — and where it fails

3 / The ledger, only where we read the cells

Table 1, white-noise selection, T = 2,520, means unless noted. Winner FWP is the family-wise false-positive rate of the selected winner at |Z| > 1.96. BIF column is the median on |Z*_WF| ≥ 0.5.

K Mean Z*ᵢₛ Mean Z*_WF ΔZ BIF^stab_0.5 Winner FWP %
1 0.81 0.77 0.04 0.65 5.3
10 1.90 0.81 1.09 1.80 40.2
100 2.79 0.82 1.97 2.60 99.9
1000 3.48 0.80 2.68 3.33 100.0

Walk-forward means stay near 0.8, which the text compares to the half-normal expectation √(2/π) ≈ 0.80. In-sample means climb. That is the picture. The K = 5, 50, 200, 500 rows exist in the paper and are omitted here only for length; they are monotone in the same direction. The 5.12 sentence is the conflict noted above, not an extra row.

Table 5, white noise, K = 400, boosting, N = 1,000 (means; BIF median on the subset):

Workflow K̂_eff Mean Z*ᵢₛ Mean Z*_WF ΔZ IS fail % WF fail %
FeatureMining 400.00 3.21 0.84 2.37 100.0 6.6
Hyperparameter tuning 2.93 2.56 0.79 1.77 88.2 4.7
TrendFollowing 5.86 2.19 0.78 1.41 64.2 5.0

Elastic-net tuning, same K = 400 white-noise control (Table 8): K̂_eff 1.727, mean |Z*ᵢₛ| 1.817, mean |Z*_WF| 0.808, mean ΔZ 1.009, IS fail 35.8%, WF fail 5.4%. Hundreds of λ values compressed; inflation shrank and did not disappear.

Two market demonstrations, gross of costs where the paper says so:

4 / How a reader could tear it down next

Not an order. A next attempt on one frozen workflow:

  1. Write the workflow down before any null seed: features, scaler (expanding or full-sample — full-sample is the leak the taxonomy names), label horizon, execution lag, search space K, selection metric, IS/WF cut. If the scaler is full-sample, Stage 1 should be able to falsify you. If it cannot, your null is not stressing the leak you think it is.
  2. Stage 1 on A–E with your own draws, not with a copied Z cutoff. Include the bounce placebo (Environment C) if the book trades close-to-close on bid/ask contaminated prices. Keep a parameter slice you did not tune on (the paper’s Θ_audit idea) or you will overfit the audit.
  3. Stage 2 ledger is ΔZ and K_eff of the thing you trade, meaning the position after the threshold, not the raw score. Table 6 is the reason. Report BIF^stab_0.5 only inside this study, and do not compare it with a paper that used another τ.
  4. Do not stop at a pass. A survivor still needs costs, a causal outer segment (Li’s trinity inequality, this slate’s Learn draft; nested CV), and a multiplicity adjustment aimed at a performance statistic (DSR). This audit’s Z is a HAC t on mean return. It is not a Sharpe deflator. They answer different questions; run the one your claim requires.
  5. If you fail only in the sparse corner of a positive control you actually have, record the activation rate. The paper’s 4.8% cell is a warning that “falsified” can mean “the walk-forward window almost never saw the state,” which is a different bug from lookahead.

5 / Verdict

The protocol is a gate plus a ledger, in that order. Induced nulls (zero mean, and one bounce placebo that is not zero mean in the observed price) ask whether the pipeline can hallucinate a walk-forward winner. ΔZ asks how far the in-sample max ran ahead of a disjoint evaluation, with K_eff so that a correlated grid is not billed as K independent bets. BIF is the unstable ratio the paper tells you not to carry across studies. None of this is a strategy, and a pass is not a ship. The FF25 case keeping its alpha, and the volatility case showing essentially no random-CV inflation, are there so the audit is not a machine that always says no. The white-noise tables are there so a K = 100 “significant” in-sample winner is not a surprise. Next attempt: your pipeline, your seeds, ΔZ on positions, then costs and DSR — still untested on this site.

Source

Numbers

Every statistic in the tables and case studies above is printed in the paper and attributed to the paper’s Monte Carlo or to the paper’s market demonstration. In particular: τ = 0.5; Bonferroni 1 − α/5 across five environments; ΔZ flag at the 0.99 quantile; Table 1 T = 2520, N = 1000; the 5.12 vs 3.33 BIF conflict at K = 1000; Table 4 error 91.5% at the one-cluster boundary; Table 5 and Table 8 at K = 400; detection 4.8% and 100% at the cited (φ, θ) corners; SPY alpha 0.39% (p 0.931), GARCH-null N = 30; FF25 Z 3.153 / alpha Z 2.708 / BIF 1.16 / 419 months; VIR 1.006. No Quant for Free rerun. summary.json not yet — no clean-room run.

Educational analysis, not investment advice. A methodology note of a publicly documented source — not a recommendation to trade or avoid any instrument, script, parameter set, or workflow. This page does not report a Quant for Free clean-room backtest and does not attach a summary.json. Figures in the text are the source’s, labeled as such. See the full disclaimer.