Tapeab.io
BETA
Home/Library/Testing a strategy
Testing a strategy5 min read

How to Evaluate a Backtest Before Trusting the Results

Ben Ghabili · Published

A smooth equity curve shows what a simulation produced. It does not tell you whether the rules were chosen with hindsight, whether the trades could have been executed, or whether the result survives outside the development sample.

Evaluating a backtest means examining its evidence and assumptions before treating its return as meaningful.

What does a backtest actually establish?

A backtest simulates how specified rules would have performed on historical data under stated assumptions. It can help investigate a strategy, but it does not establish that the same returns were earned or will occur in future trading.

Separate the strategy rules from the test machinery. The rules specify what to buy, when to enter and exit, and how positions are sized. The machinery determines data availability, simulated fills, costs, cash treatment and accounting.

A result can be reproducible yet unrealistic. Re-running the same flawed fill assumption exactly does not make that assumption valid.

Which data problems can make a backtest misleading?

A credible backtest must use information available at the simulated decision time and a clearly defined investment universe. Otherwise, the simulation can benefit from knowledge the trader would not have had.

Three questions are especially useful:

  • Was the information known then? A financial result released after the simulated entry cannot legitimately explain an earlier decision. Revised data and final classifications also require timing checks.
  • Which securities were eligible then? Using today's surviving companies to represent a historical universe can omit failures and former members. State the membership policy.
  • Could the fill happen? A rule based on a completed closing price needs an execution assumption consistent with when that price becomes available. Seeing a daily high or low does not prove an order filled at a convenient price.

These are general validity checks, not instructions for building a proprietary research system. If a report cannot explain them, the appropriate conclusion is that its evidence remains incomplete.

Does out-of-sample testing prove a strategy works?

No. Out-of-sample testing evaluates rules on data not used to develop those rules, but it is not a guarantee against overfitting or failure. The credibility of the separation depends on whether the test genuinely remained untouched and how many alternatives were tried.

If a researcher checks a holdout, adjusts the strategy after seeing the result, and checks again, that period has influenced development. Calling the final result “out of sample” does not restore the original separation.

Ask for the development and test dates, the rules frozen before testing, the number of trials or revisions where recorded, and any subsequent changes. Where trades or labels overlap across a time boundary, the split also needs to address that dependence.

A chronological split can preserve the direction of time. It still cannot make a short period representative of every future environment. Evidence across periods helps, but repeatedly searching for favourable periods creates another selection problem.

How can costs erase a positive result?

Costs reduce what a trader retains. Commissions, bid/ask spread, slippage and relevant financing or borrow costs should be modelled consistently with the instrument and execution process.

Consider 100 invented trades, each with the same fixed initial planned risk of $100, called 1R. There are 55 winners averaging +1R and 45 losers averaging −1R, before costs. These are hypothetical outcomes, not an actual backtest.

Gross result = 55 × 1R − 45 × 1R = +10R.

Gross average = 10R ÷ 100 = +0.10R per trade.

Now assume each completed trade has a flat all-in execution cost of 0.12R, including both entry and exit. This is an invented simplification, not a universal trading-cost estimate.

MeasureResult
Gross total+10R
Total costs: 100 × 0.12R12R
Net total−2R
Net average per trade−0.02R
Net result with fixed $100 risk per trade−$200

The arithmetic converts a positive gross average into a negative net average. It does not describe a compounded portfolio return, and no conclusion about statistical significance follows from these summary inputs.

The flat cost above is an assumption, not a measured cost for a real trade. Ask what evidence supports the assumed costs for the selected instrument and execution process. A credible report should explain both the base assumptions and what happens when less favourable assumptions are used.

Can strong trade results produce a weak portfolio?

Yes. Results for individual signals do not automatically translate into a portfolio that can hold them all. Simultaneous positions, capital limits, selection rules and correlated exposure affect the deployable result.

Ask what happens when more signals arrive than the account can accept. A rule that selects only the most successful historical candidates uses hindsight. A feasible rule must choose using information available at that time.

Also examine whether reported returns assume reinvestment, leverage, idle cash treatment or unlimited position size. Compare any benchmark over the same dates and on a compatible return basis. A less-invested strategy and a fully invested index have different exposure; the comparison needs that context.

Which metrics should a backtest report include?

Return needs a path and an explanation. Depending on the strategy, useful evidence includes drawdown, time below a previous peak, trade distribution, exposure, turnover, concentration and period-by-period results.

Averages can conceal dependence on a few exceptional trades. Win rate says nothing about profitability without gain and loss sizes and costs. A historical maximum drawdown is an observation, not a ceiling on future losses.

There is no single metric or minimum trade count that proves an edge. More observations can help, but highly overlapping trades are not equivalent to independent trials.

A practical backtest review checklist

QuestionEvidence to request
What was tested?Exact frozen rules, universe and date range
What information was available?Data timestamps, release timing and membership policy
How were trades executed?Order/fill assumptions, liquidity and price conventions
What did execution cost?Cost components, units and sensitivity checks
How was hindsight controlled?Development/test separation and revision history
Could the portfolio hold the trades?Capital, sizing, concurrent positions and selection rules
What drove the result?Distribution, difficult periods and concentration
What remains uncertain?Missing evidence and limits, not just a performance chart

You can apply this checklist to a report on Tapelab's research page. Evaluate the evidence in the individual report rather than assuming that a published result passes every check.

The useful conclusion may be “promising under these assumptions”, “invalid because of this data error”, or “not yet established”. A backtest earns trust through its documented boundaries, not through the attractiveness of its best number.

Nothing here is investment advice.

More from the Library
All guides