Methodology

Why you can take
these numbers seriously

Anyone can show a beautiful backtest. The question is how it was produced. This is the complete protocol — including where it can go wrong and what is built in to prevent that.

The protocol

🧱

1 · Out-of-sample firewall

Parameters are chosen on a bounded time window (for MR4H: 2023–2026). All data outside it is "blind": it never influenced the choice and serves as an independent test. In-sample and OOS numbers are always reported separately — the OOS number is the real number.

🧊

2 · Block robustness over headline numbers

Every configuration is judged per time block (e.g. 2012–16 / 2017–21 / 2022–26). A config that loses one block is out — however pretty the total. In-sample winners look good by definition; blocks reveal whether the edge survives regimes.

🧪

3 · Two independent implementations

The full logic exists twice: as a TradingView strategy (Pine) and as a separate Python port. Every logic change is verified in the port first, and only then built in Pine. Two implementations on independent data showing the same picture — only then does it count. This also catches TradingView artefacts such as repainting and optimistic fills.

🏷️

4 · Costs baked in

Spread and commission are included in every reported result (~1.1 pip for MR4H; Darwinex conditions for the OK System). No gross numbers.

📏

5 · Measured in R, not in currency

Every trade risks a fixed percentage; results are reported in risk units (R). That makes numbers comparable across tickers, periods and account sizes — and prevents the classic trick where one lucky oversized position makes the track record.

🔬

6 · Data honesty

Known data artefacts are reported, not swept away: in 2012–2013 the FX feed is missing Friday-evening bars (so a small number of trades artificially ran over the weekend), and the first DAX measurement proved contaminated because the time-exit fell outside trading hours. Both are documented — the second led to DAX's definitive rejection.

🗑️

7 · Failures get published

DAX: rejected (PF 1.017 after a clean remeasurement). The untouched-zone requirement: rejected (strangles half the trades). Extra risk on gold: rejected (return/DD ratio too poor). Counter-trade safety: regime-sensitive, off. Whoever only shows winners has something to hide.

📓

8 · Forward test as the final exam

Since 28 July 2026 both systems run live on Darwinex. Every trade — including the ones where I break my own rules — is in the journal, MT5 ticket number included. The track record is registered independently by Darwinex, not self-reported.

What is deliberately not claimed

Just as important as the evidence itself.

No PF-1.4-across-everything

No OK System configuration reaches PF 1.4 across the full 14.5 years. The edge is thin long-term and stronger in the current regime. That's stated plainly — and it's the reason for modest risk per trade.

No money promises

Backtest returns ("+227%", "+40R") are hypothetical, hindsight-computed outcomes at fixed risk percentages. They predict nothing. The forward test first has to prove for months that live does what the test promised.

No hidden discretion

Where I deviate from my own rules, it's logged as a deviation — see trade #1 in the journal, which was profitable yet is registered as a rule break, including what the result would have been with the correct stop.

Verify it yourself

The journal is live and counts from trade #1. Follow along and judge in a few months whether the forward numbers live up to the backtest.

To the forward journal