Vortex Capital Group
← Trading Insights

The 62-Session Mirage: The Sharpe 2.8 Strategy That Was Never There

Market StructureRisk Management

TL;DR - This is a research failure, published in full. We set out to build the strategy a quant fund actually runs: Avellaneda-Lee statistical arbitrage on 46 liquid US large caps - PCA eigenportfolios re-estimated nightly, each stock decomposed into factors plus an idiosyncratic residual, the residual traded against itself. 268 sessions, 5-minute full-tape (SIP) bars, every loading and beta fitted on a trailing window ending at the prior close. A three-month validation returned +11.43 bps a session at Sharpe 2.80, and we very nearly wrote that article. The full sample returns -6.09 bps (t = -1.53), with the two halves at opposite signs (-12.03 and +7.14). There is no edge here; the interesting part is how convincingly there appeared to be one. Five mechanisms, each measured: (1) the short window - 62 sessions manufactures a Sharpe of 2.8 from a strategy whose true mean is negative; (2) the sampling clock - the single-name version loses 17.77 bps at t = -2.69 on 5-minute bars but decays to -8.12, -4.55, then +5.25 at 10, 15 and 30 minutes, because it was bid-ask bounce; (3) the specification grid - 7 of 40 cells clear |t| >= 2, and not one of twenty specifications clears it in both halves with the same sign; (4) the bookkeeping error - a one-line indexing bug we caught only because it produced the number we wanted; (5) the lead that died on more data. Our one surviving result - a rate-shock effect worth +7.51 bps a shock-day (t = 2.29) - we flagged as needing more history. So we got it: 10.5 years, 2,658 sessions, 2,941 shocks, back to 2016. The 9.4 years we had never examined reverse the sign: -1.20 bps, t = -2.84 in sigma units, negative in 8 of 11 years and all 12 parameter settings - while the discovery window cleared |t| >= 2 in exactly one of those twelve, the one we reported. The "volatile rates" rescue fails too: the three most volatile years are all negative and the best year has the lowest rate vol. And the quiet-window control we called a clean zero is -0.54 bps (t = -2.86) over the full history - the same sign as the shock windows. There was never a switch. The finding is the discipline, not the strategy.

Most published trading research has a survivorship problem that has nothing to do with delisted stocks. Strategies that work get written up. Strategies that don't get quietly dropped, and the failure - which is where nearly all the information is - never leaves the desk. The result is a literature, and a social-media ecosystem, in which every backtest works, and a generation of traders who have never seen what a null result looks like from the inside.

So here is one, in full, with the numbers.

What we built

We wanted the real thing, not a moving-average crossover with a quant vocabulary. The design is the one that has been standard in equity statistical arbitrage since Avellaneda and Lee: stop thinking of a stock as a price, and start thinking of it as a bundle of factor exposures plus a residual.

Each night, on a trailing 20-session window of 5-minute returns across 46 liquid US large caps - mega-cap tech, semis, financials, healthcare, energy, staples, discretionary - we take the correlation matrix, extract its top three eigenvectors, and turn them into eigenportfolios: statistical factors the market itself defines, rather than sector labels we impose. Every name gets a beta to each factor and a residual volatility, all from that same trailing window. Then, during the session, we accumulate each name's idiosyncratic move since the open - its actual return minus what its factor exposures say it should have done - and scale it by its own residual vol. That number, the s-score, is the cleanest available answer to "how far has this stock moved for its own reasons today." At a decision time we go long the most negative, short the most positive, hedge the factor exposure out, and hold.

Two disciplines matter more than any parameter, and both are places where this kind of study usually dies quietly:

Everything is fitted on data that ends at the prior close. If you estimate the eigenvectors or betas on the same session you are trading, ordinary least squares mechanically drags the fitted residuals toward zero across the window, and the residual will "mean-revert" beautifully because you built the reversion in by construction. It is the single most common way an intraday stat-arb backtest manufactures a result. When we deliberately ran the cheating version to size the artifact, it returned -4.86 bps - in this particular specification the lookahead is not what saves the strategy, which is worth knowing, but the discipline is non-negotiable regardless.

Entries happen at the next bar's open, never at the close of the bar that generated the signal, and every figure is net of a 2 bps round-trip - conservative for names this liquid, and a cost both sides of every comparison pay.

The data is Alpaca's full consolidated tape (SIP), verified as such rather than assumed: a single 5-minute AAPL bar carries 62,554 trades on the SIP feed against 2,466 on the IEX feed. When research says "full tape," it should have checked.

Trap one: the window that flatters

We ran a three-month validation before committing to the full pipeline. It returned +11.43 bps a session, a 58% win rate, and an annualised Sharpe of 2.80. On the strength of that we started drafting an article about a working intraday alpha.

Then we ran the other ten months.

Cumulative net P&L of the intraday factor-residual reversion book - PCA eigenportfolios on 46 liquid US large caps, all loadings and betas fitted on a trailing window ending at the prior close, entry at the next bar's open, net of 2 bps round-trip. 268 sessions, Jul 2025 - Jul 2026, 5-minute full-tape (SIP) bars. The full sample returns -6.09 bps a session (t = -1.53). The final 62 sessions, shaded, return +11.43 bps a session at an annualised Sharpe of 2.80 - the exact window a three-month validation would have measured, and the reason we nearly published the opposite of this article. Compiled from public market data; VCG Research.

The shaded region is the probe window. It is not cherry-picked in the usual sense - it was chosen in advance as "the most recent quarter," which is exactly the window a practitioner would pick to check whether something still works. And inside it the strategy is genuinely excellent. It is also a fragment of a curve that spends the year going down, craters through the winter, and ends at -1,632 bps cumulative. The full-sample average is -6.09 bps a session at t = -1.53; the first two-thirds return -12.03, the last third +7.14. Opposite signs - and the earlier half is significantly negative (t = -2.55) while the recent half is not significantly anything (t = +0.99). Whichever one you had run into first, it would have told you something confident and wrong.

This is the most important chart in the article, because it is not a story about a bad strategy. It is a story about a sample size. Sixty-two sessions of a strategy whose true expectancy is around -6 bps will, a meaningful fraction of the time, produce a Sharpe near 3. A quarter is not a validation. It is a coin flip with a narrative attached, and it is the length of nearly every backtest screenshot you have ever been shown.

Trap two: significance you can tune with a clock

The one number in the whole study that cleared conventional significance was the single-name version - the trade a solo day trader could actually place, fading the day's most extreme residual with one ticket. It loses 17.77 bps a trade at t = -2.69. A real, significant, reliably negative edge is still a finding: it says the instinctive trade is a money-loser.

Except it isn't a finding either.

The single-name version - fade the day's most extreme factor residual at 11:00 ET, hold to the close - measured on four sampling clocks over the same 268-269 sessions and the same 46 names. On 5-minute bars it loses 17.77 bps a trade at t = -2.69. Coarsen the identical trade to 10, 15 and 30-minute bars and the effect decays to -8.12, -4.55 and then flips to +5.25, with no t-stat clearing 1.3. An effect that lives only on the fastest clock is bid-ask bounce, not economics. Compiled from public market data; VCG Research.

Same trade, same names, same sessions, same entry rule. The only thing that changes is the length of the bar we sample. At 5 minutes: -17.77, t = -2.69. At 10: -8.12. At 15: -4.55. At 30 it changes sign to +5.25, and none of the coarser clocks produce a t-stat past 1.3.

An economic effect does not care what bar length you view it through. Something that evaporates as you coarsen the sampling is living in the microstructure - bid-ask bounce, the mechanical negative autocorrelation of transaction prices bouncing between bid and offer. It is the same lesson as the desk's off-radar reversion illusion, where a 73% fade rate turned out to be 35% once traded properly: the appearance of an edge and the cost of trading it were the same phenomenon measured twice. If a result only exists on the fastest clock, it is a property of the quote, not of the market.

Trap three: the grid nobody publishes

Every backtest has free parameters. Ours had two that mattered: what time you decide, and how long you hold. The convention is to report the combination that worked. Here is the whole grid instead.

The full specification grid, reported rather than filtered: daily cross-sectional information coefficient of the residual signal against forward factor-hedged returns, four decision times x five holding horizons, split in-sample (Jul 2025 - Mar 2026) and out-of-sample (Apr - Jul 2026). Each cell shows both t-stats; gold outlines mark |t| >= 2. Seven of forty cells clear the bar - more than the two independent draws from noise would give, because the cells heavily overlap - but in none of the twenty specifications do both halves of the sample clear it with the same sign, which is the test that matters. Compiled from public market data; VCG Research.

Forty cells - four decision times, five holding horizons, in-sample and out-of-sample computed separately. Seven clear |t| >= 2. Pure noise across forty independent draws would deliver about two, so seven looks like something - until you note that these cells are anything but independent. They are the same trades measured at overlapping horizons from adjacent decision times, so one lucky stretch of tape lights up a whole neighbourhood of the grid at once. The test that actually discriminates is the one the grid makes visible: in none of the twenty specifications do both halves of the sample clear the bar with the same sign. The 11:00 decision time works in-sample and dies out-of-sample. The 10:00 decision time does the reverse. The 13:00 row is negative throughout the out-of-sample period and positive throughout the in-sample one.

We also ran the horizon axis specifically because a genuine mean-reversion effect must have a characteristic decay time - if a dislocation has a 30-minute half-life, its information coefficient should be largest at short horizons and fall away smoothly. There is no such shape anywhere in the grid. The ICs are 0.00 to 0.06 and wander.

And our proposed macro gate - the idea that residual reversion should pay more on days when more of the market's motion is idiosyncratic - ran backwards. Sorting sessions by their morning idiosyncratic share, the lowest quintile returned +7.83 bps and the fourth -15.84, with a non-monotone profile and an overall correlation of -0.079. We had a mechanism, a story, and a chart planned. The data declined all three.

That null is not new to this desk, and it is worth saying so: the Huddle Index work found the same thing from the other direction - the morning correlation gauge tells you how many independent bets your book is really holding, and it does not predict what the tape does next. Two studies, two different signals built from intraday correlation structure, the same verdict. It measures your risk. It does not time your entries.

What we will not tell you

While we were running this, the financial press spent July describing a market of record-low implied correlation and the highest dispersion since 2020 - the AI-capex winners and losers pulling apart. It is a good narrative and it would have been a perfect frame for this article: correlation collapses, idiosyncratic motion explodes, residual strategies come alive.

Our own measurement says the opposite. The share of intraday variance our three-factor model leaves unexplained fell from 0.674 to 0.484 across the year - realized factor structure got stronger, not weaker. Implied correlation from the options market and realized intraday factor share are genuinely different quantities and can diverge without either being wrong. But we are not going to open an article with a macro story our own data inverts, and neither should anyone else. If the narrative and the measurement disagree, publish the measurement.

The fifth trap: the lead that died on more data

We pivoted once, to the question that actually motivated the project: does macro information diffuse into the equity cross-section slowly enough to trade? With the Fed priced for hikes rather than cuts this summer, intraday rate moves are the live macro variable in this tape.

We pre-registered the hypothesis - after a sharp move in long rates, high-rate-beta names under-react and catch up - and pre-registered the stop rule. The hypothesis was rejected: the laggards do not catch up. The names that move hardest on the rate shock keep going. That is worth knowing on its own, because buying the name that "hasn't repriced yet" is the instinctive trade and it is the wrong side.

Trading that reversal - long the over-reactors, short the under-reactors, held 30 minutes and factor-hedged - returned +7.51 bps a shock-day at t = 2.29 across 295 shocks, against what looked, on those fourteen months, like nothing at all in the quiet-rate windows of the same sessions. It was the only thing in the whole program that cleared a significance bar. We wrote it up as a lead worth pursuing with more data, not something to put risk on.

Then we got the data: 10.5 years of native 15-minute full-tape bars, back to 2016 - 2,658 sessions and 2,941 rate shocks. Because the effect was discovered on Jul 2025 - Jul 2026, everything before it is a genuine held-out sample, examined for the first time with the specification frozen. No re-tuning, same 2-sigma threshold, same 30-minute hold.

The lead that did not survive more data. The rate-shock reaction-gap book - long the names that over-react to a 2-sigma 15-minute move in long rates, short those that under-react, held 30 minutes and factor-hedged - by calendar year, over 10.5 years of native 15-minute full-tape (SIP) bars: 2,658 sessions, 2,941 shocks, 46 large caps. It was discovered in the shaded window (Jul 2025 - Jul 2026, +7.51 bps a shock-day, t = 2.29). The 9.4 years before it, never examined during the original study, return the OPPOSITE sign: -1.20 bps a shock-day, t = -1.69 raw and t = -2.84 in cross-sectional sigma units, with eight of eleven years negative and a sweep of the shock threshold (1.5-3.0 sigma) x hold (15-60 min) negative in all twelve specifications - while the discovery window itself cleared |t| >= 2 in only one of those same twelve. The amber line is TLT's own 15-minute volatility, which closes the obvious escape route: the three most volatile rate years (2020, 2022, 2023) are all negative, and 2026 - the best year in the sample - has the lowest rate vol in it. Compiled from public market data; VCG Research.

It reverses. The 9.4 held-out years return -1.20 bps a shock-day, t = -1.69, and t = -2.84 in cross-sectional sigma units - the rejection is stronger on the volatility-normalised measure, not weaker. Eight of eleven calendar years are negative; 2017 is significantly negative on its own (t = -2.56). The pre-registered rule was "same sign at |t| >= 2 or it's dead." It came back opposite sign and significant. Dead.

And it is not one unlucky setting. Sweeping the two free parameters - the shock threshold at 1.5, 2.0, 2.5 and 3.0 sigma, and the hold at 15, 30 and 60 minutes - gives twelve specifications. The held-out period is negative in all twelve, six of them at |t| >= 2, every one of those six negative. Now apply the article's own trap-three test to the discovery window: across those same twelve specifications, exactly one ever cleared |t| >= 2 - the one we reported. We had, without noticing, picked the single cell in a twelve-cell grid that looked significant, which is precisely the error we spent a whole section warning about.

Two further details matter for how much you should trust this reversal. First, the rebuild is not a different method quietly producing a different answer: run on the original 14-month window, the new native-bar pipeline reproduces the original result almost exactly (+7.51 against the +6.42 we first measured on bars derived from 5-minute data). The two data paths agree. What changed is the sample, not the machinery. Second, the escape hatch does not open. The natural rescue is "the effect needs a high-rate-volatility regime, and that only arrived recently." The table refutes it: the three most volatile years for TLT - 2020, 2022 and 2023 - are all negative, and 2026, the best year in the sample, has the lowest 15-minute rate volatility of any year in it. If anything the relationship runs the wrong way.

And the correction we owe our own earlier draft: we described the quiet-rate control as a large-sample zero, and said that was what made the shock number meaningful. On 10.5 years and 25,482 quiet windows it is not zero - it is -0.54 bps, t = -2.86, the same sign as the shock windows. Shock windows and quiet windows are doing the same mild thing. There was never a switch to find.

This is the whole article happening in miniature, prospectively, with the ink still wet: an effect that cleared t = 2 on 295 events, in a sample chosen in advance, with a control group and clustered errors - and it was still the 62-session mirage wearing a better suit. The only thing that caught it was refusing to stop at the number we liked.

How the desk uses it

  • A quarter is not a validation. If your evidence for a strategy is three months, you have measured noise with a narrative attached - we produced Sharpe 2.80 out of a strategy whose real expectancy is negative, and we did it without trying. Demand a sample long enough to contain at least one regime you did not like, and split it before you look.
  • Coarsen the clock before you believe the t-stat. Re-run the identical trade on bars two, three and six times longer. A real effect survives; a microstructure artifact decays and flips. This one test would have saved us from publishing a "significant" -17.77 bps result that was bid-ask bounce.
  • Report the grid, not the cell. Before believing any parameter combination, compute all of them and count how many clear the bar in both halves of the sample with the same sign. Ours: zero out of twenty. If you only ever see the winning cell, you are looking at a selection, not a result.
  • Build the control sample - and give it the same history as the test. A conditional result means nothing until you have measured the unconditional case it is meant to differ from. Our control looked like a clean zero over fourteen months, which is exactly what made the rate-shock number seem meaningful; over 10.5 years and 25,482 windows it is -0.54 bps at t = -2.86, the same sign as the shock windows. The conditionality was an artifact of a short control, not a feature of the market.
  • Fit on data that ends before the trade. Loadings, betas, volatilities, thresholds - all of it on a trailing window ending at the prior close. Anything fitted on the session you are trading will reward you with a beautiful backtest and no money.
  • Distrust the number you wanted. Our indexing bug survived several readings because it produced a result we were hoping for. The errors that hurt are never the ones that break the code; they are the ones that improve the answer. When a result gets better, audit harder than when it gets worse.
  • When you call something "a lead needing more data," go and get the data. That phrase is usually where research goes to die politely - the promising result gets filed, repeated, and eventually cited as if it were established. Ours cleared t = 2 on a pre-registered sample with a control group, and 9.4 additional years reversed its sign. If a result is worth flagging, it is worth falsifying.

This is what most research looks like from inside a desk, and publishing only the wins is how the industry ends up with a library of strategies that stop working the moment real money touches them. The strategy in this article does not work. The process that established that it does not work is the actual deliverable - and it is the same process that, when something does survive it, tells you the edge is real.

Trade with the desk behind the research

Vortex Capital Group gives qualified traders DMA via Sterling Trader Pro, multi-vendor HTB locates, smart and dark-pool routing, and an 80%+ monthly profit share.

Apply to Trade