Vortex Capital Group
← Trading Insights

The Comb in the Fourth Decimal: What 157 Million Internalised Prints Actually Encode

Market StructureExecution Intelligence

TL;DR - There is a widely used trick for reading retail order flow off the public tape. Exchanges may not quote in increments finer than a penny, so any print with digits past the second decimal was internalised by a wholesaler, and the standard method infers the customer's side from where inside the cent it landed: just under a round penny means a buy improved off the offer, just over means a sell improved off the bid. We ran it properly. Across 157,484,769 regular-way prints in 31 US names over 15 sessions, 52.2% of prints are off-exchange, exchange prints land on the penny grid 92.64% of the time and at the half-cent midpoint another 7.26% - 0.10% everywhere else combined - while off-exchange prints sit on the grid only 34.70% of the time. The digits are real, and they are exclusive to internalisation. But they do not scatter: they land on a comb of roughly six teeth about a seventh of a cent apart, at remainders of about 0.15, 0.29, 0.43, 0.58, 0.72 and 0.86 of a cent. That comb is a distribution over the size of price improvement, and both sides of the market share it: the buy menu is 0.14, 0.28 and 0.42 cents, the sell menu is 0.15, 0.29 and 0.43 - the same six sizes offset by exactly one hundredth of a cent, which is a rounding convention, not a behaviour. The mass on the two sides is not equal. Tooth for tooth it runs 2.79x, 4.23x and 2.32x in favour of the side the rule labels buys, and 29.73% of off-exchange prints sit below a round penny against 19.37% above in 15 of 15 sessions. The decisive part is what that ratio is made of: across names it runs from 0.62x in Procter & Gamble to 6.99x in Coca-Cola, and it tracks each stock's median quoted spread (log-log r = -0.39, t = -2.29). A tight-spread stock is improved in small increments, lands above the half-cent, and reads as bought; a wide-spread stock is improved in large ones and reads as sold. Coca-Cola reads as seven-to-one bought on every session in the sample, and Procter & Gamble reads as net sold, which is a statement about their spreads rather than their customers. We then built the obvious trade anyway and tested it honestly: fading a one-standard-deviation 30-minute move earns 3.59 bp, restricting to moves retail seemed to agree with earns 7.98 bp, and the quantity actually being claimed - agreed minus fought - is 9.37 bp with a 95% interval of -2.16 to +20.90. Split by date it returned 18.53 bp at t = 3.20 over the first seven sessions and -0.16 bp at t = -0.03 over the next eight, which is the failure mode we took apart at length in The 62-Session Mirage and will not re-teach here. Nine other formulations, including the horizon the academic result actually lives at, are equally flat. The digits are worth reading. The direction inferred from them is not.

There is a number hiding in the fourth decimal place of the US tape, and for the last few years a lot of people have been trading on it.

The idea is elegant. Under Reg NMS Rule 612, no venue may display, rank or accept an order priced in an increment finer than a cent for stocks above a dollar. The rule binds quotations. It does not bind executions. So when a wholesaler internalises a retail order and gives it a fraction of a cent of price improvement, the resulting print carries digits that no exchange quote could have produced.

That makes a sub-penny print a receipt. It says: this trade did not happen on an exchange, it happened inside somebody's book.

And then comes the inference that made the technique famous. A marketable retail buy gets filled just under the offer, so its price sits a hair below a round penny. A marketable retail sell gets filled just over the bid, so its price sits a hair above one. Read the remainder, recover the side, aggregate, and you have a real-time picture of what retail is doing that costs nothing but parsing.

We wanted to know whether that picture is what it claims to be.

First, the receipt is genuine

Where a print lands inside one cent, across 157,484,769 regular-way prints in 31 US names, 2026-07-10 to 2026-07-30. Reg NMS Rule 612 binds quotations and orders, not executions, so an exchange print has almost nowhere to go: 92.64% land exactly on the penny and 7.26% at the half-cent midpoint, leaving 0.10% everywhere else combined. Off-exchange prints, which a wholesaler may price at any increment, land on the grid only 34.70% of the time; 49.10% sit somewhere in between. Those sub-penny digits exist only because the trade was internalised, which is what makes them readable at all. Consolidated SIP tape; VCG Research.

Before testing any inference we checked the premise, because everything rests on it.

Across 157,484,769 regular-way prints in 31 names over 15 sessions - Nasdaq and NYSE listings, mega caps through to sub-$200M-a-day names, $1.72 trillion of traded value - 52.2% of prints were reported off-exchange to the FINRA trade reporting facility, and those prints carried 46.4% of the dollars.

The split by price granularity is stark:

Where the print landed inside one centExchange printsOff-exchange prints
Exactly on the penny92.64%34.70%
At the half-cent (midpoint of a penny spread)7.26%16.19%
Anywhere else0.10%49.10%

An exchange print has essentially two places it can be. An internalised print has a hundred. That asymmetry is the whole basis of the technique, and it is as clean as advertised.

One more premise worth checking, since a 2024 amendment to Rule 612 would move some tick-constrained names to a half-penny quoting grid and would change all of this arithmetic: across 6,369 NBBO snapshots taken at every half-hour boundary of every session, we observed zero half-cent quoted spreads. In this sample the quote grid is whole pennies. If that changes for a name you are studying, the mapping below has to be re-derived from scratch.

The digits do not scatter. They comb.

Here is where the story stops being about retail.

If price improvement were a negotiated, continuous thing, the remainders would smear across the cent. They do not. They land on a comb of about six teeth, spaced roughly one seventh of a cent apart, at remainders of approximately 0.15, 0.29, 0.43, 0.58, 0.72 and 0.86 of a cent. Four of those six teeth are individually larger than any other non-grid remainder in the sample:

RemainderShare of off-exchange printsReads as, under the standard rule
0.586.01%discarded
0.725.41%buy, improved 0.28c
0.862.62%buy, improved 0.14c
0.432.59%discarded
0.291.28%sell, improved 0.29c
0.150.94%sell, improved 0.15c

Note what happened to the two biggest teeth. The standard rule classifies a remainder below 0.40 as a sell and above 0.60 as a buy, and throws away the middle as ambiguous. That excluded middle holds 28.30% of all off-exchange prints, and it contains the largest single tooth in the entire distribution. The method discards more prints than it assigns to either of its two most populated buckets.

That is a defect, but it is a survivable one. The next thing is not.

The comb keeps its shape and changes its position. In Coca-Cola, quoted a penny wide, almost the whole comb sits in the half the textbook rule labels buys: 62.4% of its off-exchange prints land there, and the name reads as bought on every session in the sample. In Procter & Gamble, quoted five cents wide, the same comb sits on the other side of the line and the same rule reads the name as sold. Apple, at three cents, straddles it. The ordering follows the quoted spread across all 31 names. Bars below 0.28% of prints are suppressed so the comb is legible; gold and cyan mark the six teeth, grey marks the penny and the half-cent. Consolidated SIP tape; VCG Research.

Put four names side by side on the same axis and the comb does not change shape. It changes position.

In Coca-Cola, quoted a penny wide, the improvements are small and the comb piles up hard against the top of the cent: 29.8% of its off-exchange prints sit at a remainder of 0.86 and 18.3% at 0.72. In Procter & Gamble, quoted five cents wide, the improvements are larger and the identical comb sits at the bottom: 5.8% at 0.15, 5.8% at 0.29, 4.4% at 0.43. Apple, at three cents, straddles the middle almost symmetrically.

The remainder is not a clean side flag. It encodes the side and the size of the improvement jointly, and you cannot recover one without assuming the other: a remainder of 0.72 means "0.28 cents below a round penny", which is equally describable as "0.72 cents above the penny beneath it." The standard rule resolves that by cutting the axis at half a cent, which is correct only while improvements stay under half a cent. Above that they invert, and a generously improved buy is booked as a sell.

In this sample the teeth sit at 0.14 to 0.43 cents, so the inversion is not what is doing the damage. The damage is what the comb's position does to the aggregate. In Coca-Cola almost the whole comb sits in the half labelled buys, so the name reads as bought no matter what its customers did. In Procter & Gamble almost the whole comb sits in the half labelled sells.

The mirror test

The mirror test. Both sides of the market share one improvement menu, offset by exactly one hundredth of a cent: buys are improved 0.14, 0.28 and 0.42 cents, sells 0.15, 0.29 and 0.43. That single-unit shift is the smallest step the price grid allows and is the pricing convention showing through. Compared tooth to true tooth the masses are lopsided towards the half the rule labels buys - 2.79x, 4.23x and 2.32x, and 2.28x across all six teeth taken with their neighbours. Below, each name's measured buy/sell mass ratio against its median quoted spread: the tighter the quote, the more the name reads as bought. Log-log correlation -0.39, t = -2.29, across 31 names and 82,273,630 off-exchange prints. Consolidated SIP tape; VCG Research.

There is a check that goes straight at the assumption, and it needs no model.

If the digits encode direction, the two sides of the market are two views of the same population. An improvement of a given size should show up on both. A buy improved by 0.28 cents lands at remainder 0.72; a sell improved by the same amount lands at 0.28. Whatever makes 0.28 cents a natural increment in a stock should make both of those common.

The first result was a surprise, and it is worth stating precisely because it is easy to get wrong. The menus do mirror - but offset by exactly one hundredth of a cent. Buys are improved by 0.14, 0.28 and 0.42 cents. Sells are improved by 0.15, 0.29 and 0.43. That single-unit shift is the smallest step the price grid allows, and it is the signature of a rounding rule applied in the customer's favour on each side. It is a pricing convention showing through, not a preference.

Compare the teeth to their true partners and the masses are still lopsided:

ImprovementBuy side lands atShareSell side lands atShareRatio
0.14 / 0.15 cents0.862.623%0.150.940%2.79x
0.28 / 0.29 cents0.725.406%0.291.277%4.23x
0.42 / 0.43 cents0.586.006%0.432.593%2.32x
All six teeth, with neighbours16.44%7.22%2.28x

Pooled over every remainder, 29.73% of off-exchange prints sit strictly below a round penny against 19.37% strictly above: a 1.53x skew, present in 15 of 15 sessions, ranging from 1.17x to 1.92x.

Taken alone, a two-to-four-times count skew towards buys is arguable. Retail does buy more often than it sells when the buying arrives in small dollar-denominated pieces, and the ticket sizes in this sample are consistent with that: the average identified buy was $4,708 across 14.5 million prints, the average identified sell $8,208 across 8.9 million.

What is harder to explain away is the cross-section.

The measured buy-to-sell skew is not a constant of retail behaviour. It runs from 0.62x in Procter & Gamble to 6.99x in Coca-Cola, and it tracks the name's median quoted spread: log-log correlation -0.39 across 31 names, t = -2.29. The tighter the quote, the more the name reads as bought.

Coca-Cola reads as seven-to-one net bought on every session in the sample. Procter & Gamble reads as net sold on most of them. Both are large, boring, widely held US dividend stocks, and one of them is quoted a penny wide while the other is quoted five. You can construct a story in which those two customer bases really do behave that differently. You cannot construct one in which it is a coincidence that the ordering follows the spread.

That is the state we are actually in, and it should be stated plainly rather than dressed up. The sub-penny digits carry the wholesaler's pricing convention and the customer's intention at the same time, and public data does not let you subtract one from the other. The comb's one-hundredth-of-a-cent offset is the convention showing through. The spread correlation is the convention showing through at the level of the aggregate. Whatever is left over may well be real demand - and the practical consequence is the same either way: an imbalance whose level you cannot compare between two stocks is not a measurement you can trade.

We built the trade anyway

A structural argument is not a result. So we built the most defensible version of the trade that the signal implies and tested it the way we would test our own.

The hypothesis is the good one, not the lazy one. Nobody sensible believes retail direction predicts returns on its own. The interesting claim is about counterparty: a 30-minute move that retail is participating in is a move without institutional sponsorship, and should give back; a move retail is fighting has someone real on the other side, and should hold. So: at every half-hour boundary, measure the move just completed, measure the direction of internalised flow during it, and fade the move only when the two agree.

The discipline matters more than the rule:

  • Every boundary is marked at the NBBO midpoint, not the last trade, so no result can be bid-ask bounce. A bin that ends on a retail sell prints at the bid, and the next print bouncing back to the offer would manufacture exactly the return we are looking for.
  • Every return is in excess of the cross-sectional median move of the panel over the identical window, so nothing here is a disguised market bet.
  • The comparison is against fading every qualifying move regardless of flow. Short-term reversal exists; the signal has to beat it, not ride it.
The trade built on the signal, with 95% intervals. Fading a one-standard-deviation 30-minute move earns 3.59 bp on average; restricting to moves retail appeared to agree with earns 7.98 bp. The quantity being claimed is the difference between agreeing and fighting, 9.37 bp, whose interval runs from -2.16 to +20.90 and therefore contains zero. Split by date, the rule returned 18.53 bp with a t of 3.20 over the first seven sessions and -0.16 bp with a t of -0.03 over the next eight. All returns are NBBO midpoint to midpoint and in excess of the panel's own median move over the identical window. 864 qualifying moves. VCG Research.
nMean95% intervalt
Fade every big move (the control)8643.59 bp-2.12 to 9.301.23
Retail agreed with the move4597.98 bp0.73 to 15.242.16
Retail fought the move405-1.39 bp-10.35 to 7.57-0.30
The signal: agreed minus fought8649.37 bp-2.16 to 20.901.59

The second row is the one that would have got published. On its own it clears two standard errors, and 7.98 bp against a 2-to-4 bp round trip is a business.

The fourth row is the one that matters, because the claim is not "fading works" - it is "knowing the counterparty improves fading." That quantity is 9.37 bp and its interval contains zero.

And then the check that ends the discussion. Split the sample by date - the same test that took apart a much larger piece of research on this desk four days ago:

nMeant
First 7 sessions, retail agreed20018.53 bp3.20
Next 8 sessions, retail agreed259-0.16 bp-0.03

We watched this happen in real time. At nine sessions the effect was 12.89 bp at t = 2.95, and it survived a matched-move-size control in every bucket. Two sessions later it was gone. Nothing about the method changed; the sample simply got longer.

Everything else we tested

A single null is a coincidence. Here is the whole set, all pre-specified, all on the same panel:

FormulationResult
Fade conditioned on retail agreement, 30 min9.37 bp, interval -2.16 to +20.90; flips sign out of sample
Matched inside five buckets of move sizeDifferences of 8.5, 11.7, 10.3, 5.3, 14.6 bp - every t below 1.2
Cross-sectional long/short on residual imbalancet = 1.25 gross, against a 12.8 bp round trip
Longer horizons: 60 minutes, and to the closeConsistent sign, t = 1.36 and 1.15
Morning flow (09:30-11:00) into the afternoonRaw and residualised versions have opposite signs
Retail participation as a trend-versus-chop filterTop minus bottom quintile 0.07 bp, t = 0.02
Total off-exchange share, and midpoint-print shareNon-monotonic, no surviving cut
Overnight and next-session, the horizon the published result lives atNull; raw and residualised disagree
Tight-improvement variant (only unambiguous remainders)Same magnitude, same absence of significance
Per-name demeaned imbalance, to remove the convention biasSlightly negative, t = -0.74

With 459 trades and a per-trade standard deviation of 79 bp, this sample rules out any true edge above roughly 15 bp per trade. That is not a proof of zero. But a day-trading edge has to clear costs, capacity and attention, and an effect small enough to hide behind this bound is not one you would build a desk around.

What the tape does tell you

The measurement is not worthless. It is just not a direction signal. Three things in it are large, stable and worth internalising.

Retail is a print-count phenomenon, not a dollar phenomenon. Identified retail flow is 8.20% of dollar volume but 15.5% of prints, and in Coca-Cola it is 38.2% of prints and 5.67% of dollars. The tape is dominated by tiny tickets: 72.0% of all prints in the sample are odd lots, reaching 84.5% in Nvidia and 83.9% in Johnson & Johnson. Any statistic you compute per-trade rather than per-dollar is mostly measuring people buying $300 of something.

The size of the flow varies enormously by name and is knowable in advance. It runs from 25.6% of dollar volume in SoFi and 24.0% in Marathon down to 3.3% in Carvana. If you trade a name where a quarter of the dollars are internalised retail, your resting orders are competing with a wholesaler's inventory, not with the lit book you are looking at.

Participation has a shape through the day. It runs 7.92% of dollars in the opening half hour, climbs to 10.70% around midday as institutional volume thins out, and eases into the close. The middle of the session is when the internalised share of the tape is highest, which is the opposite of where most people assume the "real" flow sits.

How the desk uses it

  • Run the mirror test on any tape-derived signal before trading it. The question generalises: if this quantity means what I think it means, what symmetry must the data have, and what must it be uncorrelated with? Then go and look. A number that moves with a property of the instrument rather than of the people trading it is measuring the plumbing.
  • Treat a public identification method as a hypothesis, not a data feed. The sub-penny rule is a good idea that a live market has quietly grown around. Somebody else's pricing convention is now inside your factor.
  • Compare against the right null. The control for "fade when retail agrees" is not zero. It is "fade regardless." That single comparison removed most of what looked like a result here.
  • Keep the parts of the measurement that are large. Off-exchange share, retail participation by name, and the odd-lot composition of the tape are all stable, all computable from the same parse, and all change how you route and size. The direction inference is the only part that failed.

The rest of the discipline - splitting by date, marking to the midpoint, counting how many specifications you tried - is the subject of The 62-Session Mirage, and the two studies failed in the same place for the same reasons.

The takeaway

The sub-penny digits are one of the most genuinely informative artefacts in the US tape. They exist because a rule about quotations does not bind executions, they are exclusive to internalised flow, and they are readable by anyone with the trades feed and an afternoon.

What they encode is the execution: which venue, and how much price improvement it carried. Direction was always an inference layered on top of that, and the inference turns out to be entangled with the pricing convention it is trying to see through - visible in a comb offset by exactly one hundredth of a cent, and in an imbalance whose level across names is predicted by the quoted spread.

The trade built on it looked like a discovery for seven sessions and was worth nothing over the next eight.

The digits are a receipt for the execution. Reading them as a record of the intention is a separate claim, and it is the one that does not survive.

At Vortex Capital Group the useful output of a study like this is not a signal. It is one more question a number has to answer before it is allowed near capital: not only is this real, but what would have to be true about the plumbing for this to be real.

Related reads

The 62-Session Mirage · Volume Shock Without Sponsorship · Liquidity Is a Clock, Not a Number · The PFOF Tax.

Joining the desk

If your first reaction to the comb was to ask what symmetry the data should have had rather than what the trade was, you think the way this desk does. The trader application takes about ten minutes; serious applicants hear back within five business days.


Methodology and sources: The panel is 31 US common stocks and ETFs (AAPL, ACHR, AFRM, AMD, CAT, COIN, CVNA, CVX, HOOD, IONQ, JNJ, JPM, KO, LCID, MARA, META, MRK, NFLX, NVDA, PG, PLTR, RIOT, RIVN, RKLB, SMCI, SOFI, SPY, TSLA, UNH, WMT, XOM) over the 15 regular sessions from 2026-07-10 to 2026-07-30, regular trading hours only, with no early-close days in the window. Trades and quotes are the full consolidated SIP tape retrieved from Alpaca Market Data v2 (/v2/stocks/trades and /v2/stocks/quotes, feed=sip). Prints are kept when the sale condition is a regular sale - a space on CTA tapes A and B, an at-sign on UTP tape C - with only the odd-lot and intermarket-sweep modifiers permitted; derivatively priced and average-price prints are excluded and counted separately. Off-exchange means exchange code D, the FINRA trade reporting facility. The sub-penny remainder is Z = 100 x (price mod $0.01), computed from the price rounded to six decimal places. Retail identification and signing follow Boehmer, Jones, Zhang and Zhang, "Tracking Retail Investor Activity", Journal of Finance 76(5), 2021: Z in (0.6, 1) is a marketable buy, Z in (0, 0.4) a marketable sell. The known signing error in that rule is documented in Barber, Huang, Jorion, Odean and Schwarz, "A (Sub)penny for Your Thoughts: Tracking Retail Investor Activity in TAQ", Journal of Finance, 2024, who placed 85,000 real trades through six brokerage accounts and report that the algorithm identifies 35% of them, mis-signs 28% of what it identifies, and yields uninformative imbalance for 30% of stocks; they propose signing against the quoted midpoint instead, which requires a full quote reconstruction this study does not attempt across the whole panel. The quotation increment rule is Rule 612 of Regulation NMS, 17 CFR 242.612, which governs displaying, ranking and accepting quotations and orders rather than executions; the 2024 amendments introduce a half-penny increment for stocks whose time-weighted average quoted spread is $0.015 or less, and we observed no half-cent quoted spreads in this sample. Every bin boundary is marked at the last valid NBBO midpoint in the seconds before the boundary; marks more than 60 seconds stale are dropped, giving 6,369 usable symbol-marks. Bin returns are midpoint to midpoint and reported in excess of the cross-sectional median return of the panel over the identical window; the move filter is a one-standard-deviation move measured per symbol over the sample. Confidence intervals are normal-approximation on the per-trade distribution; session-clustered statistics were also computed and are not materially different. Cost figures are the observed quoted spread at the entry mark. A sample of 15 sessions cannot exclude a small effect, and the bound stated in the text is the effect size this sample can exclude, not a claim of zero. No causal model is asserted: the mirror test establishes that the sub-penny distribution is inconsistent with a pure direction encoding, not the precise pricing rule that produces it, which is not recoverable from public data. Compiled from public market data - VCG Research.

#MarketMicrostructure #RetailOrderFlow #SubPenny #RegNMS #PriceImprovement #OrderFlow #QuantResearch #Backtesting #DayTrading #PropTrading #VortexCapitalGroup

Trade with the desk behind the research

Vortex Capital Group gives qualified traders DMA via Sterling Trader Pro, multi-vendor HTB locates, smart and dark-pool routing, and an 80%+ monthly profit share.

Apply to Trade