NineFifteenAM

Pillar guide · updated monthly

Walk-forward testing explained, with a worked example

8 min readBy NineFifteenAM

Walk-forward testing re-picks a setting on past data and scores it on the next unseen block. A simulated 6-window example shows the gap it exposes.

Short answer

Walk-forward testing splits your history into consecutive windows. In each one you choose a setting using only the older in-sample part, then score it once on the newer out-of-sample part that it never saw, and slide forward. In my simulated example, the in-sample choices averaged +0.21R per trade across six windows, while the same choices averaged +0.06R out-of-sample. The gap is the point: it measures how much of a result was fitting rather than a repeatable edge.

In the simulated example below, six walk-forward windows picked settings that averaged +0.21R per trade in-sample and +0.06R on the unseen data that followed. Same rule, same data, about a quarter of the apparent edge left. That shrinkage is what walk-forward testing is built to show you.

Everything in this guide uses simulated data. The numbers are illustrative. They are not from my trading and not from the market, and they say nothing about how any real strategy behaves.

What walk-forward testing is

A normal backtest picks the best setting on all your history and reports how well it did on that same history. That answer is flattering by construction, because you chose the setting by looking at it. I wrote about this in how to backtest an intraday strategy without fooling yourself.

Walk-forward testing fixes the ordering of events so it resembles real life:

  1. Take an older block of data (in-sample) and choose your setting there.
  2. Freeze that setting. Score it on the next block (out-of-sample), which it has never seen.
  3. Move both blocks forward in time and repeat.
  4. Report only the out-of-sample results, stitched together.

You are asking, again and again, "if I had picked this on the data I had then, how would it have done on what came next?"

Anchored vs rolling windows

There are two ways to move the in-sample block forward.

Anchored (expanding) Rolling (fixed length)
In-sample start Always the first date Moves forward each step
In-sample length Grows each step Constant
Data per fit More and more Same each time
Reacts to change Slowly, old data keeps its weight Faster, old data drops out
Risk Dead conditions still shape the choice Fewer trades per fit, noisier choices

Neither is correct in general. If I believe the behaviour I am testing is fairly stable, anchored uses the data best. If I suspect it drifts, rolling is more honest about it, at the price of fewer trades in each fit. The example below uses rolling windows.

The worked example

The rules of the simulation

Since the signal is noise, no value of k should do better than any other. Any apparent difference is luck, which makes this a clean test of what walk-forward reveals.

The code

import numpy as np

SEED = 915
rng = np.random.default_rng(SEED)
N_DAYS = 480
signal = rng.normal(size=N_DAYS)              # toy "signal strength", pure noise
win = rng.random(N_DAYS) < 0.42               # true win chance 42% for every day
R = np.where(win, 1.5, -1.0)                  # win = +1.5R, loss = -1R
THRESHOLDS = [0.0, 0.5, 1.0, 1.5]             # toy parameter: trade only if signal >= k
IS_LEN, OOS_LEN, MIN_TRADES = 120, 60, 15

def stats(lo, hi, k):
    m = signal[lo:hi] >= k
    r = R[lo:hi][m]
    n = len(r)
    if n == 0: return 0, 0.0, 0.0, 0.0
    losses = r[r < 0]
    return n, (r > 0).mean(), r.mean(), (losses.mean() if len(losses) else 0.0)

oos_all = []
for w in range(6):
    s = w * OOS_LEN
    i0, i1, o1 = s, s + IS_LEN, s + IS_LEN + OOS_LEN
    cands = [(stats(i0, i1, k)[2], k) for k in THRESHOLDS if stats(i0, i1, k)[0] >= MIN_TRADES]
    k = max(cands)[1]                         # choose k using in-sample days only
    a = stats(i0, i1, k)                      # in-sample result
    b = stats(i1, o1, k)                      # out-of-sample result, k frozen
    oos_all.append(R[i1:o1][signal[i1:o1] >= k])

The snippet shows the core loop; the table and the single-split comparison below come from a longer version of the same script that adds the printing. I ran it on this machine, and the output below is pasted from that run.

The results

Window Chosen k In-sample trades In-sample win rate In-sample avg R Out-of-sample trades Out-of-sample win rate Out-of-sample avg R Avg loss (R)
1 1.0 17 0.59 +0.47 8 0.38 -0.06 -1.00
2 1.0 15 0.47 +0.17 7 0.57 +0.43 -1.00
3 0.0 49 0.49 +0.22 31 0.35 -0.11 -1.00
4 0.0 57 0.44 +0.10 36 0.44 +0.11 -1.00
5 0.0 67 0.40 +0.01 38 0.50 +0.25 -1.00
6 1.0 25 0.52 +0.30 10 0.30 -0.25 -1.00

Average loss is -1.00R in every row because the simulation uses a fixed -1R loss. In a real log it would vary, and that is worth tracking.

Summary of the run:

Measure Value
Mean in-sample average R across 6 windows +0.21
Mean out-of-sample average R across 6 windows +0.06
Pooled out-of-sample: trades, win rate, average R 130 trades, 0.43, +0.08
True average R in the simulation +0.05

What the table shows

First, the in-sample numbers are always better than the out-of-sample ones on average (+0.21R against +0.06R), even though the signal does nothing. That is selection at work: I picked the best-looking setting each time, and the best-looking one is partly luck.

Second, look at the windows where the chosen k was 1.0. The in-sample trade counts were 17, 15 and 25, which is thin, and the out-of-sample counts were 8, 7 and 10. Their out-of-sample results were -0.06R, +0.43R and -0.25R. With that few trades, one or two wins moves the average a lot. The +0.43R in window 2 is not a discovery, and the -0.25R in window 6 is not a failure. Both are noise.

Third, the pooled out-of-sample figure (130 trades, +0.08R) landed close to the true +0.05R. Pooling many small windows is more informative than reading any one of them.

The single-split alternative

Instead of six windows, suppose I had used one split: the first 240 days in-sample, the last 240 out-of-sample. The same selection rule picked k = 1.0.

Trades Win rate Average R
In-sample (days 1 to 240) 32 0.53 +0.33
Out-of-sample (days 241 to 480) 48 0.40 -0.01

A single split says "+0.33R in-sample, about zero out-of-sample". Walk-forward said "+0.21R in-sample, +0.06R out-of-sample" and showed that the choice flipped between 1.0 and 0.0 from window to window. Neither is the truth, which is +0.05R. The walk-forward version is steadier because it averages over several blocks, and the instability of the chosen k is itself a warning. If the best setting changes every window, the setting probably isn't measuring anything.

Sizing the windows

There is no formula I can honestly give you. These are trade-offs I reason through:

I count trades, not days. Whatever the calendar length, I want enough trades in each block that one win or loss doesn't swing the average. I haven't measured an ideal ratio for real data, so I'd rather say that than hand you a rule of thumb.

Common mistakes

Peeking. Looking at an out-of-sample block, then changing the rules, then testing again. Once you have seen that data and responded to it, it is in-sample. The same applies if you choose window lengths because they make the final numbers look good.

Re-tuning until it passes. If the walk-forward result disappoints and you adjust the strategy and run it again, every rerun is another lottery ticket. I covered how fast that adds up in my post on opening range breakout: count your tries and treat the best as the luckiest.

Too few trades. The example shows it: windows with 7 to 10 out-of-sample trades ranged from -0.25R to +0.43R. Decide in advance how many trades a window needs before you read anything into it. Also remember that trades on the same day are related, so trade counts overstate how much independent evidence you have.

Ignoring costs. The simulation has none. Real tests should charge every trade; see what one trade really costs.

Reading averages without the shape. Average R alone hides the mix of win rate and payoff. Profit factor, payoff ratio and expectancy explains how those fit together.

Treating walk-forward as proof. It is a better filter than a single backtest, not a guarantee. The next real step after it is running the unchanged rules forward on new days, which paper trading can do cheaply.

A short checklist

  1. Write the rules and the list of settings before you look at results.
  2. Choose window lengths in advance, in trades.
  3. Choose settings on in-sample data only, and freeze them.
  4. Report the pooled out-of-sample result and the window-by-window spread.
  5. Watch how much the chosen setting changes between windows.
  6. Do not edit the rules after seeing the out-of-sample numbers.

Sources

This guide is about testing methods, using simulated data, for information only. It is not investment advice or a recommendation to trade any strategy. I am not registered with SEBI as an investment adviser or research analyst.

Questions people ask me

What is walk-forward testing?

It is a way of testing a trading idea where you repeatedly choose a setting on one block of past data, then measure it on the block that follows, then move both blocks forward in time. Only the out-of-sample blocks count as the result, so every number you report comes from trades the setting was not tuned on.

What is the difference between anchored and rolling walk-forward?

In an anchored test the in-sample window always starts at the same first date and grows longer each step. In a rolling test the in-sample window keeps a fixed length and its start date moves forward, so old data drops out. Anchored uses all history; rolling adapts faster to change but sees less data per fit.

How long should the in-sample and out-of-sample windows be?

There is no standard answer. The in-sample window needs enough trades to compare settings, and the out-of-sample window needs enough trades to say anything at all. A common starting point is an in-sample window a few times longer than the out-of-sample one, but I haven't measured which ratio is best, and it depends on how often your strategy trades.

Is walk-forward testing better than a single train/test split?

It gives you several out-of-sample checks instead of one, so a single lucky or unlucky block matters less. It does not remove overfitting: if you keep changing the rules until the walk-forward result looks good, you have turned it back into an in-sample test.

How many trades do I need in each out-of-sample window?

More than you probably have. In my simulated example, windows with 7 to 10 trades swung between -0.25R and +0.43R per trade with no real change in the underlying data. Treat any window with a handful of trades as noise and look at the pooled total instead.

walk-forward testingbacktestingout-of-sampleoverfittingpythonrolling window
N

NineFifteenAM

One trader building an options bot for Indian index markets since early 2026. I write down how it is built, what broke, and what it cost — no tips, no calls, no returns.

Related

29 Sept 2026
How to backtest an intraday strategy without fooling yourselfThe three ways a backtest lies: lookahead bias, ignored costs and overfitting. With pandas code, a coin-flip simulation, and a 68% backtest that won 20% live.
Guides
30 Sept 2026
Streaming live prices with the Kite Connect websocket in PythonA runnable Python script for Zerodha's Kite Connect websocket: subscribe, pick a mode, read a tick field by field, handle reconnects and shut down cleanly.
Guides
29 Sept 2026
Profit factor, payoff ratio and expectancy: a worked exampleWhat win rate, payoff ratio, profit factor and expectancy mean, how they connect, and 20 made-up trades worked through in R, plus where small samples mislead.
Guides