Four limits. None of them is a reason not to backtest; all of them are reasons to read one carefully.
| The limit | What it means |
|---|---|
| The past is one sample | It happened once. A strategy can be right about the mechanism and still meet a decade that does not suit it |
| Fills are assumed | You did not really trade. Every fill is a modelled price and the assumption has to be stated |
| You knew how it ended | You chose what to test knowing which years were good. That is a bias no amount of care removes entirely |
| You would not have held on | The test sat through the drawdown without flinching. You are not a test |
The most dangerous backtest is the one that works. It is the point at which people stop asking questions.
More than most publish. A few dozen tells you almost nothing about a strategy that wins most of the time and loses rarely — the losses are the part you have not sampled.
No. A backtest replays history quickly and can cover decades. Paper trading is slow and forward-looking, and it tests your behaviour as much as the rules.
It can show a strategy would have worked, on that data, under those assumptions. That is worth a great deal and it is not proof.
Because trying twenty things and showing the best one produces an impressive number from pure noise. The count is what lets a reader discount it properly.
Everything on this site comes from backtests, and none has published a result yet.
Two problems in our own data are holding them up, and both are recorded rather than worked around: a point-in-time question about how each trade was judged, and not yet being able to say which wins were quoted prices and which were modelled ones.