Automated backtests answer one question well: would these rules have made money on this data. They are silent on the question that decides your account — whether you can follow the rules when the outcome is unknown.
Manual backtesting on bar-by-bar replay is slower, and that is the point.
Write the rules down first
Before the first bar, the rules have to be specific enough that a stranger could apply them: entry condition, invalidation, target or exit rule, position size. "Buy the pullback" is not a rule. "Buy the first higher low after a break of the prior day's high, stop below that low" is.
If you cannot write it down, there is nothing to test — you will improvise, and improvisation cannot be measured.
Run a sample you did not choose
Pick a window by date, not by how the chart looks. Choosing periods where your setup obviously worked is the most common way a backtest lies, and it is easy to do without noticing.
A hundred occurrences is a reasonable first sample. Fewer than thirty tells you almost nothing.
Log every occurrence, including the ones you skipped
Skipped trades are data. If the rules fired forty times and you took twelve, you are not testing the strategy — you are testing a discretionary filter you have not written down. Either add the filter to the rules or take the trades.
Read the distribution, not the total
A profitable total built on one outlier is not an edge. Look at the spread of outcomes, the worst run of consecutive losses, and the size of the largest drawdown relative to what you would actually sit through.
Backcandle computes these per session automatically — R-multiples per trade, the longest losing streak, max drawdown and the equity curve — so the arithmetic is not what takes the time. The honest execution is.
Related: how to structure a replay session covers the constraints worth fixing before the first bar, and paper trading vs replay covers when each is the better tool.