Quant · Execution

Backtests Do Not Fill Orders

The same gross backtest is a good strategy, a marginal one, or a loss depending only on how often it trades and what it trades. Under a stated cost model, 400 basis points of gross alpha breaks even at 25 times turnover in large caps and 8 in small ones.

Issue date
Last revised
Net alpha and break-even hurdle under a stated cost modelLeft panel: simulated net annual alpha in basis points against annual two-way turnover from zero to 24 times, for three liquidity regimes, starting from a gross alpha of 400 basis points. The large-cap line stays positive across the range; the small-cap line crosses zero at about 8.4 times turnover. Right panel: the gross alpha required to break even rises linearly with turnover in each regime, crossing the 400 basis point gross assumption at the same points.Simulated: where a gross backtest stops payingsquare-root impact, 5% participation-500-2500250NET ALPHA, BP PER YEAR0x4x8x12x16x20x24xANNUAL TURNOVERNet of costs, gross alpha 400 bp02505007501000REQUIRED GROSS ALPHA, BP0x4x8x12x16x20x24xANNUAL TURNOVERBreak-even hurdle400 bp grossLarge cap 16 bp/turnMid cap 27 bp/turnSmall cap 47 bp/turnA 400 bp gross alpha breaks even at 25x turnover in large caps, 8x in small caps.The cost model is an assumption. The sensitivity to it is the point.
Net alpha and break-even hurdle under a stated cost modelLeft panel: simulated net annual alpha in basis points against annual two-way turnover from zero to 24 times, for three liquidity regimes, starting from a gross alpha of 400 basis points. The large-cap line stays positive across the range; the small-cap line crosses zero at about 8.4 times turnover. Right panel: the gross alpha required to break even rises linearly with turnover in each regime, crossing the 400 basis point gross assumption at the same points.Simulated cost sensitivity-500-2500250NET ALPHA, BP PER YEAR0x4x8x12x16x20x24xANNUAL TURNOVERNet of costs, gross alpha 400 bp02505007501000REQUIRED GROSS ALPHA, BP0x4x8x12x16x20x24xANNUAL TURNOVERBreak-even hurdle400 bp grossLarge cap 16 bp/turnMid cap 27 bp/turnSmall cap 47 bp/turn400 bp gross breaks even at 25x turnoverin large caps, and 8x in small caps.
Figure 1 · Where the alpha goes The same gross backtest is a good strategy, a marginal one, or a loss, depending only on how often it trades and what it trades. Nothing about the signal changes across these lines. This is a sensitivity experiment under an assumed cost model, not a measurement: the spread, volatility and impact coefficient are inputs chosen to bracket a plausible range, and a desk should substitute its own measured costs before drawing any conclusion about a real strategy. Source: Simulated. No market data is used in this figure. Notes: Cost per unit of turnover is half the quoted spread plus k times daily volatility times the square root of participation, with k = 0.60 and participation = 5 per cent of average daily volume. Regime parameters, spread and daily volatility in basis points: large cap 2 and 110, mid cap 8 and 170, small cap 25 and 260. The square-root impact form follows the standard concave specification; it is assumed here, not estimated. Financing, borrow, capacity and taxes are excluded, so these curves are optimistic.

Take a strategy that earns 400 basis points a year before costs, and ask only one further question: how often does it trade? Under the cost model in Figure 1, that same gross alpha survives 24 times annual turnover in large-cap US equities, breaks even at about 15 times in mid caps, and is already negative past roughly 8 times in small caps. Nothing about the signal changes across those three answers. The strategy is good, marginal, or a loss depending entirely on what it trades and how often.

This is arithmetic under stated assumptions rather than a measurement, and Figure 1 says so. But the arithmetic is where a research process either builds in its execution constraint or discovers it later at cost.

Section 01Paper versus reality

The cleanest name for the gap is old. In 1988 Andre Perold described the implementation shortfall: the difference between the return of a paper portfolio, whose trades are recorded at the prices observed when the decision was made, and the return of the real portfolio that had to be traded into existence. The backtest is the paper portfolio. Perold's point was that the cost of trading is not the commission line but the whole distance between decision and execution, including the part that never appears on an invoice because it takes the form of a worse price.

Marking an order at the midpoint records a price, not an execution. No counterparty is obliged to trade there, queue priority does not exist in the simulation, and the simulated order cannot move the book it is supposed to be crossing. A researcher can spend months refining a signal and thirty seconds assuming it fills at the mid. The assumption is silent and, for anything fast, enormous.

Section 02Where the edge goes

Figure 1 makes the arithmetic explicit. Cost per unit of turnover is half the quoted spread plus a market-impact term that scales with the square root of participation, the concave form associated with Almgren and Chriss and with the empirical impact literature. Under the parameters declared in the figure, that comes to roughly 16 basis points per unit of turnover in large caps, 27 in mid caps and 47 in small caps. Multiply by turnover and subtract.

The right-hand panel inverts the question into the number a research desk should actually be comparing its backtest against. A strategy turning over eight times a year in mid caps must produce about 214 basis points gross simply to reach zero. Comparing its backtested return to zero, rather than to that hurdle, is the specific error the panel is drawn to prevent.

Two properties of the arithmetic matter more than any single parameter. Every term subtracts, so the error introduced by a mid-price backtest is one-directional rather than symmetric. And the terms are largest for high-turnover, capacity-hungry strategies, which is exactly the sort a broad specification search tends to surface. The backtest does not err at random. It flatters, and it flatters hardest the strategies that most deserve suspicion.

The sensitivity to the impact assumption is worth stating rather than hiding. Holding everything else fixed and moving the impact coefficient across a plausible range from 0.3 to 1.5 moves the large-cap break-even from 48 times turnover to 11. Moving participation from 1 to 25 per cent of average daily volume moves it from 53 times to 12. A desk that has not measured its own impact does not know which of those worlds it is in, and the difference between them is the difference between a viable strategy and a fee-paying one.

A backtest-to-execution loss waterfall A waterfall chart on a vertical scale from zero to one hundred per cent of simulated alpha. The tall left bar is raw simulated alpha at one hundred. It is then reduced by a descending staircase of execution costs, each labelled with an illustrative percentage: fees minus seven, the bid-ask spread minus fourteen, slippage minus nine, market impact minus fifteen, latency minus five, and adverse selection minus nine. A gold bracket over those steps marks the total implementation shortfall at roughly fifty-nine per cent. What remains on the right is a cyan residual bar at roughly forty-one per cent, the edge a live strategy can actually keep. The percentages are illustrative, not measured. Backtest → execution: where the edge goes illustrative · per-trade edge, % of simulated alpha implementation shortfall ≈ 59% of raw alpha 100% 75 50 25 0 100 ≈41% −7 −14 −9 −15 −5 −9 Raw alpha Fees Spread Slippage Impact Latency Adverse selection Residual Illustrative, not measured. Steps are shares of simulated alpha; their sum is what a mid-price backtest ignores.
Figure 2 · A backtest-to-execution loss waterfall Raw alpha − fees − spread − slippage − impact − latency − adverse selection = residual edge The waterfall is a diagnostic decomposition, not a claim that every term can be estimated independently. Its purpose is to make implementation shortfall visible before overlapping execution effects are mistaken for residual noise.

Section 03You cannot trade your way out

A natural response is to trade more cleverly, and there is real science in doing so, but it has a hard limit. Almgren and Chriss framed optimal execution as a trade-off with no free corner. Trade quickly and you minimise the risk of the price drifting away while you wait, but you maximise impact. Trade slowly and you minimise impact but accept timing risk. They separated permanent impact, the lasting price change your trading causes, from temporary impact, the transient cost of demanding immediacy, and showed that the best available outcome is a point on an efficient frontier trading expected cost against the variance of that cost. There is no schedule that escapes the trade-off, only a considered choice of where to sit on it.

The deeper consequence concerns capacity. Because impact grows with size relative to depth, edge per dollar shrinks as a strategy scales, and there is some size at which the marginal trade consumes the marginal alpha. A signal can be real and still be uninvestable at scale. That is a statement about the order book rather than about the idea, and it is why execution belongs inside the research process rather than as a haircut applied at the end.

cost per unit of turnover, in basis points:

    c(q)  =  s/2  +  k * sigma * sqrt(q)

        s      quoted spread, bp
        sigma  daily volatility of the traded name, bp
        k      impact coefficient, dimensionless
        q      order size as a fraction of average daily volume

net alpha:        a_net  =  a_gross  -  T * c(q)
break-even:       T*     =  a_gross / c(q)

implementation_shortfall = paper_return - real_return
                         = fees + spread + slippage + impact + timing

A mid-price backtest sets every term on the right of the last line to zero. That is the optimism, stated precisely. The square-root impact form is an assumption in this article, not an estimate: no impact model is fitted to data anywhere in it.

A backtest does not err at random. It flatters, and it flatters hardest the strategies that most deserve suspicion.

The research object is therefore alpha conditional on an execution policy, not raw alpha. That means estimating cost against a declared benchmark, scaling impact with participation and depth, keeping rejected and unfilled orders in the sample, and testing stressed liquidity separately from average conditions. A strategy has survived implementation only when its residual return stays credible under those assumptions. Until then the backtest has established a forecast without establishing that the forecast can be owned.

  • Figure 1 is a simulation under an assumed cost model, not a measurement. The spread, volatility and impact coefficient are inputs chosen to bracket a plausible range for US equities, and a desk should substitute its own measured costs before drawing any conclusion about a real strategy.
  • The square-root impact specification is standard but not universal. Impact is known to depend on order duration, venue, and the information content of the flow, none of which the model carries.
  • Financing, borrow availability, short-sale constraints, taxes and capacity limits are all excluded, so the curves in Figure 1 are optimistic even on their own terms.
  • The cost is applied uniformly to turnover. Real desks pay very different costs on the parts of a trade they can be patient with and the parts they cannot, and that dispersion is not modelled.
  • Low-turnover strategies are far less exposed to any of this. The argument is sharpest for short-horizon and capacity-constrained strategies and weakest for slow ones.

This research is analysis and commentary for general information. It is not investment advice, an offer, or a solicitation, and it contains no price forecasts. Figure 1 is simulated under the parameters declared in its notes; the cited works are the basis for the cost specification, and the interpretation is the author's.

References & notes

  1. Perold, A. F. (1988). The Implementation Shortfall: Paper Versus Reality. Journal of Portfolio Management, 14(3), 4-9. The foundational statement of the return gap between a decision portfolio and its implemented counterpart.
  2. Almgren, R., and Chriss, N. (2000). Optimal Execution of Portfolio Transactions. Journal of Risk, 3(2), 5-39. Source for the expected-cost versus timing-risk frontier and for the separation of temporary from permanent impact.
  3. Almgren, R., Thum, C., Hauptmann, E., and Li, H. (2005). Direct Estimation of Equity Market Impact. Risk, 18(7), 58-62. The empirical basis for a concave, approximately square-root, relation between participation and impact, which Figure 1 assumes rather than estimates.
  4. The cost model, its parameters and the sensitivity results quoted in Section 02 are reproduced by the script in research/2025-03/ in the journal's repository. No market data is used in Figure 1.

Return to the front page