Quant for Free
Home / Tools / CSV Trade-Level Robustness Battery
Tool · trade-level metrology · not a signal

CSV Trade-Level Robustness Battery — Sign-Randomization, BCa, Sequence Shuffle

A single backtest is one ordering of one set of closed trades. Ushana Kevin Iorkumbul’s MQL5 piece, CSV Data Analysis (Part 7) (article 23125, 30 July 2026), bolts three reshuffles onto a trade-level CSV and prints one summary row per indicator name. This page is that battery as a metrology object. It is not a new entry rule, and it is not a second copy of the Monte Carlo trade-sequence guide already on this site.

The overlap is real and should be named. Module C — permute the order of the same trade P&Ls, rebuild the equity curve, read drawdown percentiles — is the same question as that guide’s fan chart. What Part 7 adds, and what this teardown is about, is the other two modules plus the joint report: a sign-randomization null aimed at Sortino, because order-shuffling cannot move an order-independent statistic, and a bias-corrected accelerated bootstrap interval around several trade-level metrics.

No Quant for Free book was run through this code. Nothing below is a site pass/fail.

1 / What the object actually is

Input contract: one row per completed trade. The battery requires Trade_Profit_USD and, for grouping, Indicator_Name. Other columns in the published schema (symbol, timeframe, entry/exit time, duration hours, entry slippage) are carried and not used by the three tests. The CSV is produced by TradeLevelExport.mqh, called from OnDeinit inside the Strategy Tester, matching exit deals to entries by position id. A demo EA in the article is a plain EMA cross whose only job is to emit rows. Appending several runs into one file, with different Indicator_Name values, is how variants share a report.

Three modules, three different nulls:

A. Sign-randomization permutation of the Sortino ratio. Sortino here is trade-level and not annualized:

Sortino = mean(trade profits) / downside deviation, with downside deviation = sqrt( mean( min(profit, 0)² ) ).

The article’s point, and it is the right point: shuffling order leaves this number unchanged, so an order-shuffle “null” is a spike at the observed value and tests nothing. The null that destroys a directional edge while keeping magnitudes is independent sign flips, +1 or −1 with equal probability, then recompute Sortino. The p-value is the share of sign-flipped Sortinos greater than or equal to the observed one (one-tailed). Default n_permutations in the published function: 10,000. Seed default 42. Fewer than 10 trades raises. Bands written in the interpreter: p < 0.01 “highly significant,” p < α “significant,” p < 0.10 “marginally significant,” otherwise not. Default α = 0.05.

B. Bootstrap confidence intervals, percentile and BCa. Draw the trade vector with replacement, same length, default 10,000 times. Metrics named in the default list: sortino, mean profit, win rate (percent of profits > 0), max drawdown (percent, from a cumulative equity), profit factor (gross win / gross loss, or infinity if there is no loss). The percentile interval is the empirical α/2 and 1 − α/2 quantiles at confidence 0.95. The BCa interval adds the usual bias-correction z₀ from the share of bootstrap replicates below the observed value, and an acceleration from the jackknife. The article says the primary reported interval is BCa, and that an i.i.d. resample is the wrong bootstrap when trade returns are strongly serial — a block bootstrap would be the extension, and it is not what this code does.

C. Monte Carlo order shuffle for drawdown. Permute Trade_Profit_USD, rebuild equity as initial equity plus the cumulative sum. Default initial equity 10,000. Record max drawdown in money and in percent of running peak. Default inside monte_carlo_equity_simulation: 10,000 shuffles; the unified battery calls it with 5,000. The equity fan is stored only when simulations ≤ 5,000. Percentiles reported: 5, 25, 50, 75, 95. observed_dd_pct_rank is the percent of simulated drawdowns that are ≤ the observed drawdown (a low rank means the historical ordering was a mild drawdown relative to other orderings of the same trades). Minimum length here is 5 trades, not 10.

run_robustness_battery loops indicator names, skips groups under 10 trades, writes per-variant charts and a Robustness_Summary.csv with: n trades, observed Sortino, permutation p and significance flag, BCa Sortino bounds, observed drawdown percent, MC P50, MC P95, observed drawdown rank. Charts are specified at a fixed width of 980 pixels and 120 dpi. Encodings tried, in order: cp1252, utf-8, latin-1.

Section 9’s decision framework — the author’s, not ours — says a variant “passes” only if all five hold: permutation p < 0.05; BCa Sortino lower bound positive; that interval “narrow relative to the observed value”; observed drawdown rank at or below the 40th percentile of the shuffle distribution; P95 drawdown inside a risk tolerance the user supplies. The article’s own examples of “narrow” versus “wide” are pedagogical spans (0.8 to 2.4 against an observed 1.6; 1.3 to 1.9), and the tolerance sketch “20% acceptable versus 35% P95” is likewise a hypothetical. They are not measurements of a named account.

2 / What it does not claim — and where it fails

The conclusion already refuses the promotion: the battery “quantifies validity for the tested period.” It does not guarantee future profitability, see regime shifts, or charge execution friction and overfitting. Passing “earns the right to subsequent out-of-sample testing.” That is the correct rank. A low sign-flip p-value is evidence against “these signs were a coin flip, given these magnitudes,” inside this sample. It is not evidence the magnitudes will recur, that the entries were not fit on the same tape, or that costs were real.

Failure modes:

Attached files, sizes as printed: TradeLevelExport.mqh 5.11 KB, RobustnessDemoEA.mq5 3.62 KB, robustness_testing.py 28.62 KB, generate_sample_trades.py 3.37 KB, zip 14.09 KB. We did not vendor that code into this site.

3 / How a reader could tear it down next

Not a signal. A next attempt on a CSV you already trust:

  1. Split the questions the way the article splits them. If the metric does not depend on order (Sortino, mean, win rate, profit factor), the null is sign-flip or a bootstrap, not a shuffle. If the metric is drawdown, the shuffle is the relevant toy, and it still is not a regime test. Running all three and requiring all five lights green is a screen the author proposed; failing one light is a lead, which the article also says, not an automatic discard.
  2. Freeze the trade list before the test. The CSV must be the trades of a rule whose parameters were not chosen on these same rows. Otherwise a tiny p-value restates the in-sample fit. Pair this battery with a disjoint walk-forward or the nested outer segment, not instead of it.
  3. Ablate the bootstrap. On a list you know is clumped (runs of losses), compare the i.i.d. BCa interval to a block bootstrap with a block length you pre-declare. If the lower bound flips sign, the published interval was the optimistic sampler.
  4. Define the zero-downside case and the profit-factor infinity case in advance. The published profit-factor branch returns infinity when gross loss is 0; an infinity in a summary CSV will poison any later rank.
  5. Do not double-count Module C. If you already read the trade-sequence guide, use Part 7’s shuffle as a cross-check of drawdown rank, and spend the new effort on the sign-flip p-value and the BCa bounds — the parts that guide does not implement.
  6. Costs stay outside. Slippage is a column and is not an input to Sortino in the published functions. Net-of-cost P&L has to be in Trade_Profit_USD already, or the test is on the wrong number.

4 / Verdict

Useful object, narrow claim. Sign-randomization is the correct permutation for an order-independent Sortino; order-shuffling that same Sortino is a null that cannot reject. BCa is a real interval with a real i.i.d. assumption the author does not hide. The sequence shuffle is the cousin of a guide this site already published, and it still only rearranges a fixed multiset. Five green lights would mean “this historical trade list is not a coin-flip of its own magnitudes, the Sortino lower bound stayed positive under an i.i.d. resample, and the observed drawdown was not a friendly ordering.” They would not mean the rule survives the next regime. That next test is still open, which is the right place to leave it.

Source

Numbers

From the article, not from a Quant for Free run: permutation and bootstrap defaults 10,000; unified battery Monte Carlo 5,000 (function default 10,000); α 0.05; confidence 0.95; seed 42; minimum trades 10 (permutation, bootstrap, battery skip) and 5 (Monte Carlo function); initial equity 10,000; drawdown percentiles 5 / 25 / 50 / 75 / 95; significance labels at 0.01, α, and 0.10; pass screen uses p < 0.05, positive BCa lower bound, drawdown rank at or below the 40th percentile; chart 980 px at 120 dpi; demo inputs EMA 12 / 26, lots 0.10, magic 770007. Pedagogical interval sketches 0.8–2.4 vs 1.3–1.9 and the 20% vs 35% tolerance line are examples in the prose, not results. File sizes as printed in the attachment table. No performance numbers taken from a live strategy CSV — the synthetic section states a qualitative pattern only. No site Sharpe.

Educational analysis, not investment advice. A methodology note of a publicly documented source — not a recommendation to trade or avoid any instrument, script, parameter set, or workflow. This page does not report a Quant for Free clean-room backtest and does not attach a summary.json. Figures in the text are the source’s, labeled as such. See the full disclaimer.