What a backtest does not say
A backtest answers one narrow question: how a fixed set of rules would have behaved on one sample of history. Most of the decisions taken on the back of it require answers to different questions entirely.
Philippe Trocellier — Founder — TP Advisory Services
This is a published outline, not a finished article. It sets out the argument and structure of a piece currently being written. The full text will replace it once complete.
Backtesting is a useful discipline and a poor argument. It is useful because it forces a strategy to be specified precisely enough to be executed. It is a poor argument because the result is a single realisation, produced by a researcher who already knew how the sample ended.
Three questions a backtest cannot answer alone
How much of this survives implementation?
Costs, borrow, capacity, rebalancing timing and the delay between signal and execution are modelling choices. Each is defensible in isolation; together they routinely account for the entire excess return.
How many strategies were tried?
The relevant statistic is not the Sharpe ratio of the surviving strategy, but the number of variants tested before it survived. Without that count, the significance of the result cannot be assessed.
Is the benchmark honest?
A benchmark selected after the fact turns a factor exposure into apparent skill. The comparison has to be fixed before the result is known, and it has to be investable.
The useful output of a backtest is not the equity curve. It is the list of conditions under which the strategy stops working.
A more defensible protocol
- Specify the rules, the universe and the benchmark before looking at results.
- Record every variant tested, including the discarded ones.
- Report the strategy's behaviour by regime and sub-period, not only in aggregate.
- State the implementation assumptions as a table, and show the result's sensitivity to each.
