Live Log in Start free

How many trades do you need before you trust a backtest?

Thirty trades tells you almost nothing and a hundred is often not enough. The sample you need depends on how big your edge is relative to the noise around it.

Published
Reading
3 min
Sections
4
By
Backcandle team
Feed
RSS →

"Test it on a hundred trades" is common advice and a reasonable start. It is not a rule. The number of trades you need depends on one ratio: how large your average result is compared to how much individual results swing.

A win rate is fuzzier than it looks

Say a strategy wins 60% of the time over your sample. The 95% range around that figure is roughly:

  • 30 trades: 42% to 78%
  • 100 trades: 50% to 70%
  • 300 trades: 55% to 65%

At thirty trades, a coin flip and a genuinely good setup produce numbers you cannot tell apart. Even at a hundred, the bottom of the range is break-even territory for many reward-to-risk ratios.

Expectancy needs even more

The win rate is only half of it. What matters is the average result per trade in R — the expectancy — and that number is noisier because it carries the size of every win and loss.

Take a setup that wins 40% of the time at +2R and loses 1R otherwise. Its expectancy is +0.2R per trade, a decent edge. The standard deviation of a single trade is about 1.47R, so:

  • after 100 trades, the 95% range for the average is about −0.09R to +0.49R — it still includes zero;
  • it takes around 216 trades before the range clears zero;
  • at 400 trades it narrows to +0.06R to +0.34R.

A smaller edge needs far more. Halve the expectancy and you need four times the trades — the sample grows with the square of the noise-to-edge ratio.

What this means in practice

  • Treat thirty trades as a smoke test. It can reveal a broken idea; it cannot confirm a good one.
  • Plan for a few hundred occurrences of a setup before sizing up on it, more if the edge is thin or the payoff is lopsided.
  • Keep the sample honest. Every occurrence the rules flagged counts, including the ones you skipped. A large sample of hand-picked trades is still a biased sample.
  • Split it. If the first half and the second half disagree badly, you have learned something the total would have hidden.

Getting to a few hundred

The obstacle is time, not arithmetic. A setup that appears twice a week takes two years of live trading to reach two hundred occurrences. On bar-by-bar replay the same sample is a matter of weeks, because the quiet hours between setups can be skipped. The case for compressed practice time makes that argument at length.

Backcandle's journal reports win rate, expectancy and profit factor per session, so the numbers above are easy to keep up with as the sample grows. For running the sample itself, see how to backtest a futures strategy without writing code.

Keep reading

More from
the team.