A backtest is a story you tell yourself about the past. Written carefully, it's a useful story. Written carelessly, it's a very convincing lie, and the lie is always in your favour, because the person who wrote it is also the person who wants it to work.
I've been on the receiving end. One of my first strategies arrived with a 68% backtest win rate and won 20% of its trades once it ran forward. This guide covers the three ways backtests mislead, with the examples I ran into, so you can check for each one before you trust a curve.
1. Lookahead bias: using tomorrow's newspaper
Lookahead bias is when the test lets a decision use data that didn't exist yet at that moment. It is the most common bug in a first backtest and the hardest to see, because nothing crashes. The results just look wonderful.
Here is a typical example on candle data. The rule is "buy when a candle closes above the day's high so far":
import pandas as pd
df = pd.read_csv("nifty_5min.csv", parse_dates=["time"]) # open, high, low, close per candle
# highest high of the day so far, not counting the current candle
df["prior_high"] = df.groupby(df["time"].dt.date)["high"].transform(lambda s: s.cummax().shift(1))
# WRONG: the signal uses this candle's close, and the trade is booked at this candle's open.
# You can't know the close when the candle is just opening.
df["signal"] = df["close"] > df["prior_high"]
df["entry_price"] = df["open"]
The signal depends on the close of candle t, but the entry is at the open of candle t. In real time the close doesn't exist until the candle is over. The fix is to act on the next candle:
# RIGHT: decide at the end of candle t, enter at the open of candle t+1
df["signal"] = df["close"] > df["prior_high"]
df["entry_price"] = df.groupby(df["time"].dt.date)["open"].shift(-1) # next candle's open, same day
trades = df[df["signal"] & df["entry_price"].notna()] # each row's entry price is the NEXT candle's open
Other places where lookahead hides:
- Indicators that repaint. Some indicators change their past values when new data arrives. In a backtest you see the final version, not what you'd have seen live.
- The day's high or low. Using "the high of the day" in a rule that runs at 10 AM quietly uses prices from 2 PM.
- Adjusted data. Prices adjusted for splits and bonuses after the fact can change levels you'd have seen at the time.
- Stops that cheat. A stop-loss that always fills at exactly the stop price, even when the candle gapped through it, uses information about the path that you wouldn't get.
A good habit is to print the data available to the strategy at one specific timestamp, and check by eye that nothing in the future is in there.
2. Costs and fills: the edge that isn't there
A backtest that ignores charges compares a strategy with a market where trading is free. It isn't.
On Zerodha, an options trade pays ₹20 brokerage per executed order, STT on the sell side, exchange transaction charges, SEBI fees, stamp duty on the buy side, and GST on the brokerage and exchange charges. For one lot on a ₹100 premium, the round trip costs close to a rupee per unit before you count slippage, and more as a share of cheaper options. The full calculation is in brokerage, STT and slippage: what one trade really costs.
Fills matter just as much. A backtest usually buys at the last traded price or the candle's close. A real order meets the other side of the order book. On a fast-moving option, that can be several points, which is often the very moment a breakout fires.
Three rules I follow now:
- Charge every trade its real costs, including both legs.
- Assume you get filled on the wrong side of the book. Buy at the ask and sell at the bid, or add a fixed slippage allowance you've chosen deliberately.
- Test how much cost the strategy can survive. Double the costs. If the edge disappears, it was never comfortable.
3. Overfitting: try enough things and one will work
This is the one that surprised me most, because it can happen with no bug at all.
Every test you run is a lottery ticket. If you try 25 versions of a strategy and pick the best, you haven't found the best strategy. You've found the luckiest of 25.
I wanted to see how bad it gets, so I ran a simulation of strategies with no edge at all. Each "strategy" is a pure coin flip: every trade wins or loses one unit with equal chance, so the true edge is exactly zero before costs. Each strategy takes 250 trades. I generated 25 of them, picked the best, and repeated that 10,000 times.
What 25 strategies with zero edge look like:
| Result | |
|---|---|
| A single coin-flip strategy finishing at +30 units or better | 3% of the time |
| The best of 25 coin-flip strategies finishing at +30 or better | 58% of the time |
| Median result of the best of 25 | +30 units |
| Median win rate of the best of 25 | 56% |
| Best of 25 finishing at +20 or better | 96% of the time |
A win rate of 56% and a profit of 30 units, from strategies with no edge at all. If I had shown you that one, you'd have been impressed.
Now take off a small cost of 0.1 units per trade, which is tiny in real life. The median best result falls from +30 to about +5. The lucky winner barely survives costs, and that's before the market changes.
I tested 25 versions of the opening range breakout. The simulation is the reason I don't trust the best of them just because it's the best of them. Published results deserve the same treatment: the US opening-range paper I built one version from didn't carry over to Nifty.
How to protect yourself
- Count your tries. Keep a list of every version you test, including the ones you dropped. The number matters more than you'd think.
- Change one thing at a time, and write down why before you run the test.
- Fewer settings. A strategy with two knobs has less room to fit noise than one with ten.
- Look at days, not just trades. Trades on the same day are related. What matters is how many separate days the result comes from.
- Prefer a plateau to a peak. If a stop of 12 works but 11 and 13 fail, you've found noise. If 10 to 14 all work, you may have found something.
The worked example: a backtest that reversed
My clearest example is the classic mechanical 15-minute opening range breakout. Its backtest covered 530 trades, with a 68% win rate and a profit factor of 7.9. That is the kind of number that makes you stop checking.
Traded forward, on 45 real trades after I picked it, only 20% were winners and the average trade lost about 13% of the option premium.
The backtest had assumed things a real market doesn't give you: stops that followed the price perfectly, and fills at exactly the price the signal wanted. That's the kind of optimism that section 1 and section 2 describe, and it was enough to turn a strong-looking result into a losing one.
It's worth being fair to the small sample here: 45 forward trades is not a lot, and a 20% win rate on 45 trades has a wide margin: roughly 11% to 34% at 95% confidence. Even the top of that range is half the backtest's 68%, so the gap is far too large to put down to chance, and the direction was the same as everything else I'd seen.
A checklist before you trust a backtest
- Write the rules down first, with every parameter, before you look at any results.
- Check for lookahead. Print what the strategy can see at one timestamp. Enter on the next candle, not the current one.
- Include real costs, both legs, and a slippage allowance.
- Count every version you tried, and treat the best as the luckiest until it proves otherwise.
- Look at the number of days, not only trades.
- Check the plateau. Nearby settings should also work.
- Forward test. Run the rules unchanged on new days, in paper mode first. My paper trading guide covers how far to trust it.
- Start small if it passes, and keep the log honest. How I analyse my trading bot's trades is the process I use.
None of this makes a backtest useless. It makes it a first filter: a way of throwing out ideas cheaply, not a way of proving one.
Sources
- Zarattini, Barbon and Aziz: A Profitable Day Trading Strategy for the U.S. Equity Market (an example of a published intraday result worth testing yourself rather than trusting)
- The simulation described above: 25 strategies of 250 fair coin-flip trades, repeated 10,000 times. It is my own and uses no market data.
This guide is about testing methods, for information only. It is not investment advice or a recommendation to trade any strategy. I am not registered with SEBI as an investment adviser or research analyst.