← Back to the blogLeer en español
Strategy research

Walk-forward: test your strategy beyond training

Separate selection from evaluation, configure Deepwick windows and understand the limits of out-of-sample results.

After comparing parameters, the best result belongs to a configuration selected precisely because it stood out on those data. To check whether the rule remains useful in another period, separate selection from evaluation. That is the purpose of a walk-forward test.

What changes from a single backtest

Training uses part of the history to select parameters. Testing uses a later part with those parameters fixed. Repeat the process across chronological windows.

The Probability of Backtest Overfitting studies the risk of selecting strategies through repeated historical tests. Walk-forward helps separate stages, but cannot remove selection bias if you keep redesigning a strategy after inspecting its test results.

A window example, not a performance promise

Suppose you have 1,000 candles and choose four blocks with 60% allocated to training. Each 250-candle block contains 150 training candles and 100 test candles:

BlockTrainingTest with fixed parameters
1Candles 1–150Candles 151–250
2Candles 251–400Candles 401–500
3Candles 501–650Candles 651–750
4Candles 751–900Candles 901–1000

This describes the block split in Deepwick's current tool. It is not a training window that automatically accumulates every previous block. The counts explain the split; they do not guarantee enough trades or adequate warm-up for every indicator.

Configure it in Deepwick

You need an account and a saved backtest with numeric inputs. If you are still in the demo, first save your strategy through the dashboard workflow. Open your backtests, enter a report and choose walk-forward. The segments link is a different analysis, not the same selection-and-evaluation procedure.

Choose an input under Parameter. For a channel length, for example, try From 10, To 30, Step 5. That gives five values: 10, 15, 20, 25 and 30. This is an instructional range, not a recommendation for trading parameters.

Set Folds to 4 and Train ratio to 0.6 if your sample supports that split, then press Run walk-forward. The current interface sweeps one parameter at a time. If the selector is empty, return to a saved backtest containing numeric inputs. If you see fold_too_small, use fewer blocks or more history: each side needs at least 50 candles to execute. That technical minimum does not validate a strategy.

Interpret the output

Mean OOS Sharpe summarizes the test-window Sharpe ratios; it is not necessarily the Sharpe of a single concatenated equity curve. Mean OOS Return averages test returns rather than computing a compounded portfolio return. Parameter Stability measures how often consecutive windows select the same value.

A negative out-of-sample Sharpe can reflect a weak rule, the period, costs or a small sample. It does not prove overfitting by itself. A positive value does not establish a future edge either. Stable selection deserves inspection, but can be consistently unprofitable.

This tool uses its own execution and cost assumptions. Do not compare its figure with the demo as though both configurations were identical. Keep the script, sweep range and settings for each experiment so you can explain differences.

When a test stops being new

If you inspect results, change the range and repeat until they improve, you have indirectly used those test windows to select the system. Record every iteration and reserve a later period or paper-trading phase for another evaluation.

Before starting, write down what observation would lead you to reject the hypothesis. That prevents every unfavorable result becoming a reason to search for another combination. To review the initial metrics, return to how to read a backtest.

BacktestingWalk-forward