Anyone can show a beautiful backtest. The question is how it was produced. This is the complete protocol — including where it can go wrong and what is built in to prevent that.
Parameters are chosen on a bounded time window (for MR4H: 2023–2026). All data outside it is "blind": it never influenced the choice and serves as an independent test. In-sample and OOS numbers are always reported separately — the OOS number is the real number.
Every configuration is judged per time block (e.g. 2012–16 / 2017–21 / 2022–26). A config that loses one block is out — however pretty the total. In-sample winners look good by definition; blocks reveal whether the edge survives regimes.
The full logic exists twice: as a TradingView strategy (Pine) and as a separate Python port. Every logic change is verified in the port first, and only then built in Pine. Two implementations on independent data showing the same picture — only then does it count. This also catches TradingView artefacts such as repainting and optimistic fills.
Spread and commission are included in every reported result (~1.1 pip for MR4H; Darwinex conditions for the OK System). No gross numbers.
Every trade risks a fixed percentage; results are reported in risk units (R). That makes numbers comparable across tickers, periods and account sizes — and prevents the classic trick where one lucky oversized position makes the track record.
Known data artefacts are reported, not swept away: in 2012–2013 the FX feed is missing Friday-evening bars (so a small number of trades artificially ran over the weekend), and the first DAX measurement proved contaminated because the time-exit fell outside trading hours. Both are documented — the second led to DAX's definitive rejection.
DAX: rejected (PF 1.017 after a clean remeasurement). The untouched-zone requirement: rejected (strangles half the trades). Extra risk on gold: rejected (return/DD ratio too poor). Counter-trade safety: regime-sensitive, off. Whoever only shows winners has something to hide.
Since 28 July 2026 both systems run live on Darwinex. Every trade — including the ones where I break my own rules — is in the journal, MT5 ticket number included. The track record is registered independently by Darwinex, not self-reported.
Just as important as the evidence itself.
No OK System configuration reaches PF 1.4 across the full 14.5 years. The edge is thin long-term and stronger in the current regime. That's stated plainly — and it's the reason for modest risk per trade.
Backtest returns ("+227%", "+40R") are hypothetical, hindsight-computed outcomes at fixed risk percentages. They predict nothing. The forward test first has to prove for months that live does what the test promised.
Where I deviate from my own rules, it's logged as a deviation — see trade #1 in the journal, which was profitable yet is registered as a rule break, including what the result would have been with the correct stop.
The journal is live and counts from trade #1. Follow along and judge in a few months whether the forward numbers live up to the backtest.
To the forward journal