In the simulated example below, six walk-forward windows picked settings that averaged +0.21R per trade in-sample and +0.06R on the unseen data that followed. Same rule, same data, about a quarter of the apparent edge left. That shrinkage is what walk-forward testing is built to show you.
Everything in this guide uses simulated data. The numbers are illustrative. They are not from my trading and not from the market, and they say nothing about how any real strategy behaves.
What walk-forward testing is
A normal backtest picks the best setting on all your history and reports how well it did on that same history. That answer is flattering by construction, because you chose the setting by looking at it. I wrote about this in how to backtest an intraday strategy without fooling yourself.
Walk-forward testing fixes the ordering of events so it resembles real life:
- Take an older block of data (in-sample) and choose your setting there.
- Freeze that setting. Score it on the next block (out-of-sample), which it has never seen.
- Move both blocks forward in time and repeat.
- Report only the out-of-sample results, stitched together.
You are asking, again and again, "if I had picked this on the data I had then, how would it have done on what came next?"
Anchored vs rolling windows
There are two ways to move the in-sample block forward.
| Anchored (expanding) | Rolling (fixed length) | |
|---|---|---|
| In-sample start | Always the first date | Moves forward each step |
| In-sample length | Grows each step | Constant |
| Data per fit | More and more | Same each time |
| Reacts to change | Slowly, old data keeps its weight | Faster, old data drops out |
| Risk | Dead conditions still shape the choice | Fewer trades per fit, noisier choices |
Neither is correct in general. If I believe the behaviour I am testing is fairly stable, anchored uses the data best. If I suspect it drifts, rolling is more honest about it, at the price of fewer trades in each fit. The example below uses rolling windows.
The worked example
The rules of the simulation
- Seed: 915, with NumPy's
default_rng. Anyone can reproduce the numbers. - Length: 480 simulated trading days, one candidate trade per day.
- Signal: each day gets a random "signal strength" drawn from a standard normal distribution. It is pure noise and carries no information about the result.
- Outcome: each trade wins with probability 0.42 for every day, regardless of the signal. A win is +1.5R and a loss is -1R. The true average is therefore 0.42 x 1.5 - 0.58 x 1 = +0.05R per trade, a very small real edge.
- Toy parameter: a threshold k from 0.0, 0.5, 1.0 or 1.5. The rule takes a trade only on days when the signal is at least k.
- Windows: rolling, 120 days in-sample, 60 days out-of-sample, stepping forward 60 days, for 6 windows.
- Selection: in each in-sample window pick the k with the highest average R, but only among settings with at least 15 trades.
Since the signal is noise, no value of k should do better than any other. Any apparent difference is luck, which makes this a clean test of what walk-forward reveals.
The code
import numpy as np
SEED = 915
rng = np.random.default_rng(SEED)
N_DAYS = 480
signal = rng.normal(size=N_DAYS) # toy "signal strength", pure noise
win = rng.random(N_DAYS) < 0.42 # true win chance 42% for every day
R = np.where(win, 1.5, -1.0) # win = +1.5R, loss = -1R
THRESHOLDS = [0.0, 0.5, 1.0, 1.5] # toy parameter: trade only if signal >= k
IS_LEN, OOS_LEN, MIN_TRADES = 120, 60, 15
def stats(lo, hi, k):
m = signal[lo:hi] >= k
r = R[lo:hi][m]
n = len(r)
if n == 0: return 0, 0.0, 0.0, 0.0
losses = r[r < 0]
return n, (r > 0).mean(), r.mean(), (losses.mean() if len(losses) else 0.0)
oos_all = []
for w in range(6):
s = w * OOS_LEN
i0, i1, o1 = s, s + IS_LEN, s + IS_LEN + OOS_LEN
cands = [(stats(i0, i1, k)[2], k) for k in THRESHOLDS if stats(i0, i1, k)[0] >= MIN_TRADES]
k = max(cands)[1] # choose k using in-sample days only
a = stats(i0, i1, k) # in-sample result
b = stats(i1, o1, k) # out-of-sample result, k frozen
oos_all.append(R[i1:o1][signal[i1:o1] >= k])
The snippet shows the core loop; the table and the single-split comparison below come from a longer version of the same script that adds the printing. I ran it on this machine, and the output below is pasted from that run.
The results
| Window | Chosen k | In-sample trades | In-sample win rate | In-sample avg R | Out-of-sample trades | Out-of-sample win rate | Out-of-sample avg R | Avg loss (R) |
|---|---|---|---|---|---|---|---|---|
| 1 | 1.0 | 17 | 0.59 | +0.47 | 8 | 0.38 | -0.06 | -1.00 |
| 2 | 1.0 | 15 | 0.47 | +0.17 | 7 | 0.57 | +0.43 | -1.00 |
| 3 | 0.0 | 49 | 0.49 | +0.22 | 31 | 0.35 | -0.11 | -1.00 |
| 4 | 0.0 | 57 | 0.44 | +0.10 | 36 | 0.44 | +0.11 | -1.00 |
| 5 | 0.0 | 67 | 0.40 | +0.01 | 38 | 0.50 | +0.25 | -1.00 |
| 6 | 1.0 | 25 | 0.52 | +0.30 | 10 | 0.30 | -0.25 | -1.00 |
Average loss is -1.00R in every row because the simulation uses a fixed -1R loss. In a real log it would vary, and that is worth tracking.
Summary of the run:
| Measure | Value |
|---|---|
| Mean in-sample average R across 6 windows | +0.21 |
| Mean out-of-sample average R across 6 windows | +0.06 |
| Pooled out-of-sample: trades, win rate, average R | 130 trades, 0.43, +0.08 |
| True average R in the simulation | +0.05 |
What the table shows
First, the in-sample numbers are always better than the out-of-sample ones on average (+0.21R against +0.06R), even though the signal does nothing. That is selection at work: I picked the best-looking setting each time, and the best-looking one is partly luck.
Second, look at the windows where the chosen k was 1.0. The in-sample trade counts were 17, 15 and 25, which is thin, and the out-of-sample counts were 8, 7 and 10. Their out-of-sample results were -0.06R, +0.43R and -0.25R. With that few trades, one or two wins moves the average a lot. The +0.43R in window 2 is not a discovery, and the -0.25R in window 6 is not a failure. Both are noise.
Third, the pooled out-of-sample figure (130 trades, +0.08R) landed close to the true +0.05R. Pooling many small windows is more informative than reading any one of them.
The single-split alternative
Instead of six windows, suppose I had used one split: the first 240 days in-sample, the last 240 out-of-sample. The same selection rule picked k = 1.0.
| Trades | Win rate | Average R | |
|---|---|---|---|
| In-sample (days 1 to 240) | 32 | 0.53 | +0.33 |
| Out-of-sample (days 241 to 480) | 48 | 0.40 | -0.01 |
A single split says "+0.33R in-sample, about zero out-of-sample". Walk-forward said "+0.21R in-sample, +0.06R out-of-sample" and showed that the choice flipped between 1.0 and 0.0 from window to window. Neither is the truth, which is +0.05R. The walk-forward version is steadier because it averages over several blocks, and the instability of the chosen k is itself a warning. If the best setting changes every window, the setting probably isn't measuring anything.
Sizing the windows
There is no formula I can honestly give you. These are trade-offs I reason through:
- Longer in-sample windows give more trades per fit and steadier choices, but they reach further back into conditions that may no longer apply.
- Shorter in-sample windows adapt faster but, as window 1 and 2 show, fit on 15 or 17 trades, which is mostly noise.
- Longer out-of-sample windows give more trades to judge, but fewer windows overall and a longer wait before re-tuning.
- Shorter out-of-sample windows give more windows but each one tells you little.
I count trades, not days. Whatever the calendar length, I want enough trades in each block that one win or loss doesn't swing the average. I haven't measured an ideal ratio for real data, so I'd rather say that than hand you a rule of thumb.
Common mistakes
Peeking. Looking at an out-of-sample block, then changing the rules, then testing again. Once you have seen that data and responded to it, it is in-sample. The same applies if you choose window lengths because they make the final numbers look good.
Re-tuning until it passes. If the walk-forward result disappoints and you adjust the strategy and run it again, every rerun is another lottery ticket. I covered how fast that adds up in my post on opening range breakout: count your tries and treat the best as the luckiest.
Too few trades. The example shows it: windows with 7 to 10 out-of-sample trades ranged from -0.25R to +0.43R. Decide in advance how many trades a window needs before you read anything into it. Also remember that trades on the same day are related, so trade counts overstate how much independent evidence you have.
Ignoring costs. The simulation has none. Real tests should charge every trade; see what one trade really costs.
Reading averages without the shape. Average R alone hides the mix of win rate and payoff. Profit factor, payoff ratio and expectancy explains how those fit together.
Treating walk-forward as proof. It is a better filter than a single backtest, not a guarantee. The next real step after it is running the unchanged rules forward on new days, which paper trading can do cheaply.
A short checklist
- Write the rules and the list of settings before you look at results.
- Choose window lengths in advance, in trades.
- Choose settings on in-sample data only, and freeze them.
- Report the pooled out-of-sample result and the window-by-window spread.
- Watch how much the chosen setting changes between windows.
- Do not edit the rules after seeing the out-of-sample numbers.
Sources
- The simulation described above: seed 915, 480 simulated days, 6 rolling windows. It is my own, uses no market data, and is illustrative only.
This guide is about testing methods, using simulated data, for information only. It is not investment advice or a recommendation to trade any strategy. I am not registered with SEBI as an investment adviser or research analyst.