EA Settings Optimization Tested: 519 Pips In-Sample Became 4.4 Out-of-Sample

Key takeaways
- I optimised 129 parameter sets on the first half of 16,926 EURUSD 5-minute bars, then ran the winners on the second half, which they had never seen. The grid champion made +519.6 pips in-sample and +4.4 out-of-sample. That is 99.2% of the result gone.
- Across all 128 usable settings, the rank correlation between the two halves was 0.063. Knowing how a setting performed on history told me almost nothing about how it would rank on new data.
- The best in-sample setting in each family ranked 9th on average out-of-sample. One flipped sign completely: RSI 9 at level 20 went from +447.7 to -156.6 pips.
- The one setting that travelled was the slowest and least tuned: MA 5/150, +193.6 then +106.3. It was also the only one of four still standing after a single pip of spread.
Every optimiser in MetaTrader does the same thing: it tries thousands of settings on past data and hands you the one with the tallest equity curve. The question nobody in the results window answers is whether that number belongs to the strategy or to the sample. So I split my dataset in half and asked it directly.
How was the optimisation tested?

Same 16,926 EURUSD 5-minute bars, 3 June to 26 August 2026, that sit under every test in this series, so the numbers line up with my moving average crossover test and my ADX above 25 test.
- The split: bar 8,463 divides the set. In-sample is 3 June to 15 July. Out-of-sample is 15 July to 26 August. Nothing from the second half touches the choice of settings.
- The grid: 129 parameter sets across four families — 88 moving average crossovers, 16 RSI reversal rules, 9 Donchian breakout lengths and 16 Bollinger band touches. 128 produced at least 30 signals in each half and were kept.
- The measurement: move in pips after a fixed 24-bar hold, two hours, in the signal direction. No stop, no target, no money management, so the number describes the entry rule rather than a scheme layered on top of it.
- The rule: pick the highest total pips in the first half. Report what that same setting did in the second half. This is exactly what the MT5 strategy tester does when you click Start.
What happened to the settings that won on history?

| Family | Best in-sample | In-sample pips | Out-of-sample pips | Rank out-of-sample |
|---|---|---|---|---|
| Bollinger touch | 20, 2.0 SD | +519.6 | +4.4 | 6 of 15 |
| RSI reversal | RSI 9, level 20 | +447.7 | -156.6 | 9 of 16 |
| MA crossover | 5/150 | +193.6 | +106.3 | 18 of 88 |
| Donchian breakout | 100 bars | -102.0 | +100.1 | 3 of 9 |
Read the top row slowly, because it is the whole article in one line. The Bollinger 20/2.0 touch was the single best rule in a 129-set grid on the first half of the data: 227 signals, 56.4% win rate, +2.289 pips per trade. On the second half the same rule made 230 signals, won 47.0% of them and finished at +0.019 pips per trade. Same rule, same instrument, same timeframe, six weeks later. The edge was not damaged, it was never there — it belonged to the sample I picked it from.
The RSI row is worse than gone, it is inverted. A rule that made 1.972 pips a trade in June and early July lost 0.725 a trade in late July and August. If you had traded it live off that backtest you would not have suffered decay, you would have suffered the opposite of what you were promised.
Averaged across the four families, the picks went from 1.488 pips per trade in-sample to 0.343 out-of-sample. Seventy-seven percent of the performance was selection.
How much of the ranking survived the switch?
Total pips is one number and it moves for many reasons, so I also asked a cleaner question: does the order of the settings hold? If optimisation works at all, the settings that ranked highest on the first half should at least tend to rank high on the second. I measured the rank correlation between the halves for every family.
| Family | Settings | Rank correlation, half to half |
|---|---|---|
| MA crossover | 88 | +0.574 |
| Donchian breakout | 9 | +0.133 |
| Bollinger touch | 15 | +0.086 |
| RSI reversal | 16 | -0.015 |
| All settings pooled | 128 | +0.063 |
Zero would mean the backtest ranking carries no information about the future ranking whatsoever. Pooled across 128 settings I measured 0.063, which is zero with a rounding error attached. RSI came out fractionally negative, meaning the settings that looked best on history were, if anything, marginally worse afterwards.
The exception is the interesting part. Moving average crossovers held at +0.574 — not reliable, but genuinely informative. The difference is what the parameter controls. An MA pair sets the structure of the rule: how slow it is, how many trades it takes, what size of move it responds to. An RSI level sets a threshold against a distribution that shifts as soon as volatility does. Structure travels. Thresholds do not. That is the same split I found when I tested Stochastic settings on MT4, where the oversold line was the least stable input in the indicator.
How much did the optimiser gain by simply choosing?
Here is the size of the illusion. For each family I compared the winning setting to the median setting in the same family, on the same in-sample data.
| Family | Winner minus median, in-sample |
|---|---|
| Bollinger touch | +328.5 pips |
| Donchian breakout | +294.7 pips |
| RSI reversal | +228.4 pips |
| MA crossover | +194.3 pips |
Between 194 and 329 pips of apparent performance were manufactured by the act of picking the top of a list. None of it required the strategy to be good. Search a wide enough grid on a fixed sample and something will always sit at the top, for the same reason that somebody always wins a raffle.
The mirror image confirms it. The genuinely best setting on the second half, chosen with hindsight, was MA 12/20 at +201.0 pips out-of-sample — and it had made +1.7 pips in-sample, ranking nowhere. The optimiser could never have found it. Meanwhile Donchian at 100 bars was chosen out of a family where every single setting lost money in-sample, and it went on to make +100.1 out-of-sample. It was not selected for being good. It was selected for being the least bad, and then the market changed.
Does walk-forward optimisation fix it?
The standard answer to curve fitting is walk-forward: keep re-optimising on a rolling window and trade the winner on the next window. I ran it — five equal windows, four handoffs, the full 129-set grid re-searched each time.
| Optimised on | Winner | Train pips | Traded on | Live pips |
|---|---|---|---|---|
| 03 Jun – 19 Jun | RSI 21, level 35 | +248.1 | 19 Jun – 07 Jul | +14.5 |
| 19 Jun – 07 Jul | RSI 9, level 20 | +251.9 | 07 Jul – 23 Jul | +13.3 |
| 07 Jul – 23 Jul | MA 50/100 | +125.2 | 23 Jul – 10 Aug | +102.9 |
| 23 Jul – 10 Aug | MA 12/20 | +218.5 | 10 Aug – 26 Aug | -35.7 |
Three of four windows finished positive and the sequence totals +95.0 pips, which looks like a rescue until you notice that the winner changed family twice in four handoffs and that 108% of the profit came from one window. Train totals between 125 and 252 pips delivered live results between -36 and +103. Walk-forward did not remove the fitting. It made it repeat, faster.
What does the spread do to all of this?
Everything above is gross. My spread and swap cost measurement puts a realistic EURUSD 5-minute cost near one pip a round turn, so I subtracted exactly one pip per trade from every setting’s out-of-sample result.
| Out-of-sample, 128 settings | Share profitable |
|---|---|
| Before costs | 67.2% |
| After 1 pip per trade | 21.1% |
Of the four optimised picks, one survived: MA 5/150, at +106.3 pips across 89 trades, keeps +17.3 after costs. The Bollinger champion goes from +4.4 to -225.6, because 230 trades at a pip each is a bill that a 0.019 pip edge cannot pay. The walk-forward sequence goes from +95.0 to -193.0. This is why trade count matters as much as win rate, a point I made when testing stop distances: the cheapest way to lose to costs is to be right slightly, often.
One number that changes how I read any backtest
67.2% of settings were profitable out-of-sample against 52.3% in-sample. The second half of the summer was simply an easier stretch for these rules than the first. That is the quiet finding underneath everything else: the half of the data mattered more than the setting. When the regime moves, it moves every parameter set at once, which is precisely why a number tuned inside one regime cannot describe the next one.
What I would actually do with this
- Never accept an optimiser result you have not split. Hold back at least a third of the data, and look at that third before you look at anything else.
- Prefer structural parameters over threshold parameters. Lengths and speeds carried a 0.574 rank correlation. Levels carried -0.015.
- Judge the neighbourhood, not the peak. If a setting is good only because its neighbours are bad, you have found a hole in the sample, not an edge.
- Subtract the spread before you get excited. It cut the survivor rate from 67.2% to 21.1% here.
- Treat walk-forward as a test, not a cure. It exposed the instability honestly. It did not remove it.
Frequently asked questions
Does optimising an EA actually work?
It works far less than the results window implies. Across 128 settings the in-sample ranking carried a 0.063 rank correlation with the out-of-sample ranking, and the grid champion kept 0.8% of its in-sample pips. It is not useless — three of four optimised picks still beat the textbook default settings out-of-sample — but the improvement is a fraction of what the backtest shows.
What is the difference between in-sample and out-of-sample?
In-sample is the data you used to choose the settings. Out-of-sample is data the choice never touched. Any performance figure taken from in-sample data includes the benefit of hindsight; only the out-of-sample figure is an estimate of the future. Here the gap between them was 1.488 versus 0.343 pips per trade.
How much data should I hold back?
I used a straight 50/50 split, which left about six weeks and at least 30 signals per setting on each side. A third is the usual minimum. The binding constraint is trade count, not calendar length: 30 trades is the floor below which I refuse to draw a conclusion at all.
Is walk-forward optimisation better than a single split?
It is more honest, not more profitable. In this run it produced +95.0 pips gross and -193.0 after one pip of spread, with the winning strategy family changing twice in four handoffs. Its real value is diagnostic: if the winner keeps changing, you have learned that there is nothing stable to find.
What data was this run on?
16,926 EURUSD 5-minute bars, 3 June to 26 August 2026 — the same set behind my Ichimoku cloud breakout test, my Fibonacci retracement test and my news release test. The split point is bar 8,463.
Want a clean indicator to install right now?
It is my own enhanced DeMARK Trend Line indicator for MetaTrader 4 and 5. Non repaint, clean, and free.

