Quant for Free
Home / Analysis / Session VWAP Bias
Strategy teardown · Session VWAP Bias · QQQ

Does Session VWAP Bias day trading survive validation?

Zarattini–Aziz: long above session VWAP, short below, flat by the close — imbalance as a day-trade system.

It is one of the cleanest day-trading sentences in the retail literature: treat session VWAP as the imbalance line, stay long while price is above it, short while below, and be flat by the cash close. Zarattini and Aziz put that sentence on SSRN with a MATLAB backtest on QQQ that looked like a holy grail. We rebuilt the idea clean-room in Python — no MATLAB, no Pine — on Alpaca SIP, with the same cost and Deflated Sharpe Ratio gauntlet as every other teardown here.

It does not clear the bar we use for a tradeable edge after realistic friction. That is not the same as “VWAP is empty.” At zero cost on the paper’s calendar window, our 1-minute clean-room prints the same shape the authors advertised: ~17% hit rate, tens of thousands of flips, large cumulative return, single-digit-ish max drawdown. Turn on the site’s 1.5 bps of slip and the same tape goes to nearly −100%. The useful residue is where the gross imbalance story still asks a real question, and which next experiment is actually worth the trial budget.

Credit. Carlo Zarattini and Andrew Aziz, Volume Weighted Average Price (VWAP) The Holy Grail for Day Trading Systems (SSRN 4631351, November 13, 2023; also Concretum PDF). The rule we locked is the public strategy definition: regular-hours session VWAP from typical price × volume, long when the close is above VWAP, short when below, reverse when a candle closes across VWAP, no overnight. Distinct from the published NY 15-minute Opening Range Breakout teardown — that is a time-window breakout; this is a session-VWAP imbalance bias. No vendor script was copied.

1 / The pitch: above VWAP long, below short, done by 16:00

The pitch fits under a 1-minute or 5-minute chart:

Sold as an imbalance edge: institutions benchmark execution to VWAP, so price persistently above the line means net buying pressure that can continue — and vice versa. If that sentence is true as a standalone system on QQQ once you pay to flip twenty thousand times, a clean-room with 1.5 bps of friction should not need a miracle sample to show it.

Paper claims (authors’ MATLAB, not our run). On QQQ, Jan 2 2018 → Sep 28 2023, they report +671% total return, Sharpe 2.1, MDD 9.4%, ~21,967 trades, ~17% hit ratio, gain:loss ≈ 5.67, versus buy-and-hold QQQ +126% / Sharpe 0.7 / MDD ~35.6%. TQQQ variant: +8,242%, Sharpe 1.7, MDD 36.1%. Their cost assumption: $0.0005 per share commission and no slippage on a $25k start. Authors also note profits concentrate in the morning and the last hour.

Our costs (site default, documented). Fees 0. Slippage 1.5e-4 (1.0 bps spread + 0.5 bps slip) — same framing as the NY ORB teardown. Signal at the close of the bar that prints the VWAP decision; Portfolio.from_orders with targetpercent fills at that close (paper’s close-across-VWAP semantics). Optuna: 30 trials on buffer_bps 0–5, long_only on/off, skip_lunch on/off (flat 12:00–15:00 ET). Chronological 70 / 30 IS / OOS. DSR via pipeline.metrics. Tape: Alpaca SIP. Primary searchable result is QQQ 5-minute RTH. Locked default also runs on QQQ 1-minute RTH for paper-bar fidelity. SPY 5-minute is a secondary check. We do not headline Yahoo’s ~60-day intraday window.

2 / TEST A: QQQ 5-minute RTH, 30 Optuna trials, site costs

Universe: QQQ Alpaca SIP 5m, RTH only. Span: 2016-01-04 09:30 America/New_York → 2026-08-28 15:55. 208,692 bars. IS through 2023-06-16; OOS after.

Best trial: buffer_bps=5, long_only=true, skip_lunch=false.

QQQ 5m RTH · Session VWAP · 30 trials
Best params5 bp buffer, long only, no lunch skip
In-sample Sharpe (per bar)−0.002180
Out-of-sample Sharpe (per bar)−0.009871
Deflated Sharpe (DSR)≈ 0.000006
OOS return−37.91%
OOS max drawdown−38.72%
OOS trades / win rate1,809 / 29.68%

The search did not find a winner hiding in the buffer. It found the least-bad way to lose: long-only, five basis points of room before a flip, still negative in-sample, still negative out-of-sample. DSR is effectively zero against a bar we treat as about 0.95. verdict_draft.survives is false.

Default full sample (buffer 0, long/short, no lunch skip) on the same 5m tape: 20,325 trades, win rate 20.65%, return −99.03%, max drawdown −99.08%, per-bar Sharpe −0.017684. That is death by turnover under 1.5 bps — not a mysterious alpha hole.

QQQ 5m RTH Session VWAP Bias OOS equity vs buy-and-hold
QQQ Alpaca 5m SIP RTH · best Optuna trial · out-of-sample, after site costs, vs buy-and-hold.

3 / TEST B: QQQ 1-minute locked default — paper bar, honest costs

Same locked rule as the paper’s bar size. No Optuna. Full Alpaca 1m RTH span 2016-01-04 09:30 → 2026-08-28 15:59 (1,042,215 bars): 44,279 trades, win rate 13.34%, return −100.0%, max drawdown −100.0%, per-bar Sharpe −0.018836 under site slip.

On the paper calendar window only (2018-01-02 → 2023-09-28), still under site 1.5 bps slip: 23,216 trades, win rate 14.07%, return −99.19%, MDD −99.22%. Trade count is within a few percent of the authors’ 21,967. Hit rate is in the same teen band as their 17%. The frequency of the system matches. The net result, once you charge for each flip, does not.

QQQ 1m RTH Session VWAP Bias default full-sample equity
QQQ Alpaca 1m SIP RTH · locked default · full sample after site costs vs buy-and-hold.

4 / Cost sensitivity on the paper window (still our clean-room)

This is not a second headline. It is the honesty check that explains why our DSR fails while the SSRN abstract still looks like a brochure.

QQQ 5m RTH, locked default, 2018-01-02 → 2023-09-28 (buy-and-hold on the same slice: +129.14%):

Cost assumptionTradesWin rateReturnMax DDSharpe / bar
Zero cost10,81123.98%+305.70%−18.84%0.009954
~Paper fee (2e−6), 0 slip10,81123.95%+288.53%−19.25%0.009665
0.5 bps slip10,81123.06%+37.63%−28.53%0.002777
1.5 bps slip (site)10,81121.30%−84.16%−85.84%−0.011248

QQQ 1m RTH, same window (buy-and-hold +128.71%):

Cost assumptionTradesWin rateReturnMax DDSharpe / bar
Zero cost23,21616.99%+753.79%−10.92%0.006635
1.5 bps slip (site)23,21614.07%−99.19%−99.22%−0.013524

Read that table twice. Zero-cost 1m on the paper window is in the same neighborhood as the authors’ story: ~23k trades, ~17% wins, large positive return, ~11% max drawdown. We are not claiming we replicated their 671% / Sharpe 2.1 — different vendor, different fill engine, no MATLAB. We are claiming the gross object is not imaginary. We are also claiming that under the friction we use for every other day-trade teardown, the same object does not survive.

SPY 5m secondary (same Optuna space, site costs): best again buffer_bps=5, long-only; IS Sharpe −0.006098; OOS Sharpe −0.008412; DSR ≈ 0; OOS return −26.95%. Same verdict family.

SPY 5m RTH Session VWAP Bias OOS equity
SPY Alpaca 5m SIP RTH · best Optuna trial · out-of-sample after site costs.

5 / The illusions this tape is good at producing

Win rate. Paper ~17%. Our zero-cost 1m paper window 16.99%. Our site-cost OOS best trial on 5m ~29.7% on far fewer trades after the optimizer dropped the short side and added a buffer. None of those numbers is a system. Trend-day VWAP bias is supposed to lose often and win large. Win rate is the photogenic statistic; equity after costs is the adult one.

“It works in the paper.” It works in the paper’s cost world: near-zero per-share commission and assumed-zero slippage on a small account. That is a legitimate research choice to disclose. It is not the cost world we use to decide whether something is shippable next to Supertrend and ORB on this site.

Tuning as rescue. Thirty trials on 5m, every in-sample Sharpe negative. The optimizer’s best idea was “trade less and don’t short.” That is SPY/QQQ talking — a structural bid — not a discovered edge. Deflation is then almost ceremonial: you cannot deflate your way into an edge the in-sample peak already refused to show under site costs.

Bar size as hidden leverage. 1m flips more than 5m. More flips means more times you pay the spread. Zero-cost 1m looks better than zero-cost 5m on the paper window (+753% vs +305%). Site-cost 1m looks worse (−99% vs −84% on that window). The bar that makes the brochure prettier is the bar that kills you fastest once slip is real.

6 / What this does not kill (untested on purpose)

A failed DSR under site costs is not a ban on session VWAP as a question. These are the next attempts we would actually run. None of the lines below have numbers attached unless summary.json already has them.

  1. Paper-faithful costs as a labeled sensitivity, not a headline. We already ran zero / paperish / 0.5 bps / 1.5 bps on the paper window. A full Optuna + DSR path under a disclosed paperish cost would answer a different question (“does the authors’ world still hold on Alpaca?”). It would not answer “can you trade this at retail.”
  2. Time-of-day filter. Authors show morning + last-hour concentration and midday weakness. Our skip_lunch flag was in the Optuna space; under site costs the best trial did not want it. A stricter “only 09:31–11:00 and 15:00–16:00” rule is a different strategy and needs its own trial count.
  3. Stop / reverse semantics. We reverse on close across VWAP (and optional buffer). Wick-through, a fixed tick buffer, or a minimum hold are different fills.
  4. VWAP as a feature, not a system. Distance to VWAP, minutes since cross, and morning-only bias are features you can feed a different edge (ORB confirmation — the authors’ other paper — vol-target, overnight/intraday split). This run tests standalone “always in with the session VWAP.”
  5. Not QQQ cash, or not unlevered. Paper’s TQQQ result is a leverage story on top of the same rule. We did not run TQQQ. Futures (NQ) change both the cost curve and the roll.
  6. Anchored / multi-day VWAP. Out of scope. Different object.

If you take one research prompt from this piece, take the cost curve seriously: re-run the locked 1m rule with a disclosed slip grid (0, 0.25, 0.5, 1.0, 1.5 bps), publish the break-even, and only then decide whether a time-of-day filter is worth spending trials on.

7 / What you'd need to believe anyway

To treat Session VWAP Bias as a standalone, shippable QQQ system from this run, you would need to believe some combination of:

We don’t. We also don’t believe “DSR failed, delete VWAP from your brain.” Session VWAP is still where cash index liquidity is measured. That is a reason to keep measuring. It is not a reason to ship an always-in flipper at 1.5 bps.

8 / Verdict

Verdict · survives = false

Not a tradeable standalone edge on this tape under site costs. Still a real gross question.

The named Session VWAP Bias day trade, as Zarattini–Aziz describe it, was searched 30 ways on QQQ Alpaca SIP 5-minute RTH from 2016-01-04 to 2026-08-28. Best trial: 5 bp buffer, long only. In-sample Sharpe −0.002180. Out-of-sample Sharpe −0.009871. Deflated Sharpe ≈ 0. Out-of-sample return −37.91%, max drawdown −38.72%, 1,809 trades, 29.68% wins. It does not survive.

The locked 1-minute default under the same 1.5 bps slip is worse in the only way that matters: full-sample return −100%, and on the paper’s own 2018–2023 window return −99.19% with 23,216 trades and a 14.07% win rate — trade count and hit rate in the same band as the paper, net result inverted by friction.

What the zero-cost sensitivity still allows: on that same paper window, 1m clean-room return +753.79%, MDD −10.92%, win rate 16.99%, 23,216 trades. That is why this teardown is not “VWAP is fake.” It is “VWAP-as-always-in-system is a cost story.” Paper claims remain paper claims (+671% / Sharpe 2.1 / MDD 9.4% on QQQ under their commission and zero-slip assumptions).

What we would run next, and have not as a headline: a disclosed slip grid with break-even bps, a morning-plus-last-hour-only variant with its trial count in the title, and VWAP distance as a feature next to ORB rather than a standalone book. Session VWAP is still allowed to be a benchmark. It is not, on this evidence, a product at retail friction.

Check the next claim the same way

Tool
Deflated Sharpe Ratio — 30 trials, OOS Sharpe −0.009871, DSR ≈ 0
Tool
Net-vs-Gross Costs — this run’s headline used 1.5 bps slip; zero-cost / 0.5 bps sensitivities are in summary.json
Related · Teardown
NY 15-minute ORB — same authors’ family, different object (time-window breakout vs session-VWAP bias)
Related · Teardown
Supertrend — another simple rule that lives or dies with costs and the sample market
Learn · Module 5
The validation gauntlet — overfitting, OOS, DSR, and costs
Educational analysis, not investment advice. A methodology case study of a publicly popular session-VWAP bias day-trade rule — associated with Zarattini & Aziz (SSRN 4631351), reimplemented clean-room — not a recommendation to trade or avoid any strategy, parameter set, or instrument. Simulated and optimized results depend on data source, session definition, costs, position mode, fill assumption, and sample window; they do not predict future performance. Paper claims are labeled separately from clean-room numbers. Fields this run did not emit are omitted on purpose. See the full disclaimer.