Anyone can show a beautiful backtest. The question is how it was produced. This is the complete protocol — including where it can go wrong and what is built in to prevent that.
Parameters are chosen on a bounded time window (for MR4H: 2023–2026). All data outside it is "blind": it never influenced the choice and serves as an independent test. In-sample and OOS numbers are always reported separately — the OOS number is the real number.
Every configuration is judged per time block (e.g. 2012–16 / 2017–21 / 2022–26). A config that loses one block is out — however pretty the total. In-sample winners look good by definition; blocks reveal whether the edge survives regimes.
The full logic exists twice: as a TradingView strategy (Pine) and as a separate Python port. Every logic change is verified in the port first, and only then built in Pine. Two implementations on independent data showing the same picture — only then does it count. This also catches TradingView artefacts such as repainting and optimistic fills.
Spread and commission are included in every reported result (~1.1 pip for MR4H; Darwinex conditions for the OK System). No gross numbers.
Every trade risks a fixed percentage; results are reported in risk units (R). That makes numbers comparable across tickers, periods and account sizes — and prevents the classic trick where one lucky oversized position makes the track record.
Known data artefacts are reported, not swept away: in 2012–2013 the FX feed is missing Friday-evening bars (so a small number of trades artificially ran over the weekend), and the first DAX measurement proved contaminated because the time-exit fell outside trading hours. Both are documented — the second led to DAX's definitive rejection.
DAX: rejected (PF 1.017 after a clean remeasurement). The untouched-zone requirement: rejected (strangles half the trades). Extra risk on gold: rejected (return/DD ratio too poor). Counter-trade safety: regime-sensitive, off. Whoever only shows winners has something to hide.
Since 28 July 2026 both systems run live on Darwinex. Every trade — including the ones where I break my own rules — is in the journal, MT5 ticket number included. The track record is registered independently by Darwinex, not self-reported.
Just as important as the evidence itself.
No OK System configuration reaches PF 1.4 across the full 14.5 years. The edge is thin long-term and stronger in the current regime. That's stated plainly — and it's the reason for modest risk per trade.
Backtest returns ("+227%", "+40R") are hypothetical, hindsight-computed outcomes at fixed risk percentages. They predict nothing. The forward test first has to prove for months that live does what the test promised.
Where execution deviates from the rules, it's logged as a deviation — profitable trades still get registered as rule breaks, including what the result would have been by the book.
The live test runs on a fresh account with bot execution, benchmarked against the systems test. Follow along and judge in a few months whether the forward numbers live up to the backtest.
To the forward journal