Zero of twenty cells passed.
Numerroo took five simple rules that ship inside Keel, two momentum rules and three mean-reversion fades, and ran them against the same standard Assay applies elsewhere: real out-of-sample data, two timeframes, measured execution costs, and a pass line fixed before the result was known.
Across all twenty out-of-sample test cells, calculated as five rules multiplied by two timeframes and two cost treatments, not one survived. That is not an embarrassment to hide. It is the premise the engine is built on, measured on our own code and reported whether it flattered us or not.
Five rules shipped inside Keel.
The subject of this Assay is internal. Keel's architecture rests on a claim recorded in its own research ledger: simple, fittable price rules do not carry a repeatable edge once real costs are charged. That is a convenient thing to believe about your own design. Assay 003 stops believing it and measures it.
The five rules are Keel's shipped implementations, not reconstructions: a moving-average crossover and Donchian-channel breakout for momentum, plus a z-score fade, RSI fade and Donchian fade for mean reversion. The configuration was fixed in advance and never adjusted after any result was seen. Parameters were crossover 10/30, Donchian period 20, z-score lookback 20 at two standard deviations, RSI 14 with 30/70 thresholds, fixed 15/25 basis-point brackets, and the three fades restricted to the quiet 0 to 6 UTC window while momentum rules traded unrestricted.
Positive account return and benchmark outperformance.
At the frozen configuration, executed on EUR/GBP after the measured cost model, each rule needed to produce a positive whole-account out-of-sample result and beat simply buying and holding the pair over the same window under the same cost treatment.
A rule survived for a timeframe only if both legs held. Fail either leg and the claim was falsified for that cell. Every rule was read twice: once in-sample to see whether an apparent edge existed, and once out-of-sample to see whether it survived. Only the out-of-sample read could pass or fail the claim.
One frozen judge, two cost treatments.
OANDA EUR/GBP mid candles at one-hour and thirty-minute timeframes, from 2 January 2022 to 29 May 2026, split chronologically 70% in-sample and 30% out-of-sample.
All five rules and both timeframes were reported together. No cell was selected after seeing the outcomes.
About 2.06 bps round trip for momentum and 1.80 bps for fades. This was lower than the original assumption and therefore kinder to the rules.
A fixed 2.5 bps round trip, retained as a second treatment so the original internal study can be reproduced faithfully.
One long entry at the first evaluation bar and one exit at the last, charged one round trip of the same cost and sized identically to the rule.
Datasets, costs and configuration were frozen and checksummed before evaluation. Results publish identically regardless of outcome.
Every combination remains inspectable.
The controls below expose every combination. Switch timeframe and cost treatment to update the table, profit-factor chart, net-result chart and rule explorer from the sealed metrics file.
Showing H1 at measured market cost.
| Rule | In-sample PF | Out-of-sample PF | OOS net | Buy-and-hold | Verdict |
|---|---|---|---|---|---|
| MA crossover 10/30610 trades | 0.75 | 0.78 | -7.38 | +1.96 | Fails net |
| Donchian breakout 20448 trades | 0.68 | 0.73 | -6.83 | +1.96 | Fails net |
| Z-score fade25 trades | 1.15 | 0.78 | -0.30 | +1.96 | Fails net |
| RSI fade171 trades | 0.98 | 0.83 | -1.47 | +1.96 | Fails net |
| Donchian fade126 trades | 1.10 | 0.94 | -0.39 | +1.96 | Fails net |
Inspect one cell in full.
The account finished negative after cost, so the cell failed before benchmark comparison was needed.
Positive is not the same as passed.
Exactly one cell in the whole grid ended out-of-sample with a positive number. The RSI fade on the thirty-minute timeframe, at measured market cost, finished up 0.28. On a naive reading, that becomes the headline: “RSI fade is profitable out-of-sample.” It is exactly the sentence that can sell a course.
Three pre-committed checks dismantle it. First, it is the only positive cell in twenty, which is entirely consistent with statistical noise. Second, it appears only at the cheaper measured cost. At the original 2.5 bps cost it loses money. Third, the same rule and timeframe lost in-sample with a profit factor of 0.90. Even the positive out-of-sample result trailed buy-and-hold, so it fails the claim regardless.
Two different failure shapes.
The moving-average crossover and Donchian breakout sit well below a profit factor of one in both samples. They turn over hundreds of round trips and produce the largest losses. There is no measurable edge here for costs to erase.
Several fades look plausible in-sample and then fall below one out-of-sample. That shape is consistent with a rule fitted to its own history rather than a repeatable edge.
A narrow result, stated firmly.
This Assay is about five specific rules, at frozen parameters, on one pair, at two timeframes, over one window, under two cost treatments. It is not a claim that markets are unbeatable, that all systematic trading fails, or that price-based methods can never work. It says nothing about other pairs, parameters, sessions or windows, and it is not a prediction of the future.
What it does say is small and firm: run Keel's own simple price rules the way a real account pays for them, judge them out-of-sample against a bar fixed in advance, and none survived.
The receipts stay attached.
Every figure traces to a frozen, checksummed evidence capsule containing the configuration, per-cell metrics, receipt and conformance record. The configuration was pre-registered before results existed. The dataset and measured costs were frozen independently. The evaluation was run once and reconciled against the sealed metrics.
Our own rules received no special treatment.
They were frozen before testing, evaluated out-of-sample, charged real costs and reported exactly as they performed. Every cell failed. The result is not that trading is impossible. The result is that evidence must outrank intuition, including our own.
This is independent research on simulated paper execution for educational purposes. It is not financial, investment or trading advice, and nothing here is a recommendation to buy, sell or use any instrument, strategy, system or product.
Cost assumptions are operator measurements and assumptions and must be verified against current venue schedules. Real-world results vary with venue, size and conditions. Past or simulated performance does not indicate future results.