Assay 003 · Published

Keel puts its own trading rules through Assay.

Numerroo froze the five simple price rules shipped inside Keel, charged real execution costs, and judged them out-of-sample. Across twenty test cells, none survived.

Internal mechanism test EUR/GBP · H1 and M30 2022-01-02 to 2026-05-29 Frozen evidence attached
The result

Zero of twenty cells passed.

Numerroo took five simple rules that ship inside Keel, two momentum rules and three mean-reversion fades, and ran them against the same standard Assay applies elsewhere: real out-of-sample data, two timeframes, measured execution costs, and a pass line fixed before the result was known.

Across all twenty out-of-sample test cells, calculated as five rules multiplied by two timeframes and two cost treatments, not one survived. That is not an embarrassment to hide. It is the premise the engine is built on, measured on our own code and reported whether it flattered us or not.

0 / 20cells falsified after cost and benchmark testing

One cell finished slightly positive. It still lost to buy-and-hold and failed under the original study cost.

What was tested

Five rules shipped inside Keel.

The subject of this Assay is internal. Keel's architecture rests on a claim recorded in its own research ledger: simple, fittable price rules do not carry a repeatable edge once real costs are charged. That is a convenient thing to believe about your own design. Assay 003 stops believing it and measures it.

The five rules are Keel's shipped implementations, not reconstructions: a moving-average crossover and Donchian-channel breakout for momentum, plus a z-score fade, RSI fade and Donchian fade for mean reversion. The configuration was fixed in advance and never adjusted after any result was seen. Parameters were crossover 10/30, Donchian period 20, z-score lookback 20 at two standard deviations, RSI 14 with 30/70 thresholds, fixed 15/25 basis-point brackets, and the three fades restricted to the quiet 0 to 6 UTC window while momentum rules traded unrestricted.

The claim under Assay

Positive account return and benchmark outperformance.

At the frozen configuration, executed on EUR/GBP after the measured cost model, each rule needed to produce a positive whole-account out-of-sample result and beat simply buying and holding the pair over the same window under the same cost treatment.

A rule survived for a timeframe only if both legs held. Fail either leg and the claim was falsified for that cell. Every rule was read twice: once in-sample to see whether an apparent edge existed, and once out-of-sample to see whether it survived. Only the out-of-sample read could pass or fail the claim.

How it was measured

One frozen judge, two cost treatments.

Data

OANDA EUR/GBP mid candles at one-hour and thirty-minute timeframes, from 2 January 2022 to 29 May 2026, split chronologically 70% in-sample and 30% out-of-sample.

Configuration

All five rules and both timeframes were reported together. No cell was selected after seeing the outcomes.

Measured market cost

About 2.06 bps round trip for momentum and 1.80 bps for fades. This was lower than the original assumption and therefore kinder to the rules.

Original study cost

A fixed 2.5 bps round trip, retained as a second treatment so the original internal study can be reproduced faithfully.

Benchmark

One long entry at the first evaluation bar and one exit at the last, charged one round trip of the same cost and sized identically to the rule.

Publication rule

Datasets, costs and configuration were frozen and checksummed before evaluation. Results publish identically regardless of outcome.

Detailed results

Every combination remains inspectable.

The controls below expose every combination. Switch timeframe and cost treatment to update the table, profit-factor chart, net-result chart and rule explorer from the sealed metrics file.

Timeframe
Cost treatment

Showing H1 at measured market cost.

Selected Assay 003 results
RuleIn-sample PFOut-of-sample PFOOS netBuy-and-holdVerdict
MA crossover 10/30610 trades0.750.78-7.38+1.96Fails net
Donchian breakout 20448 trades0.680.73-6.83+1.96Fails net
Z-score fade25 trades1.150.78-0.30+1.96Fails net
RSI fade171 trades0.980.83-1.47+1.96Fails net
Donchian fade126 trades1.100.94-0.39+1.96Fails net
Profit factor
In-sample compared with out-of-sampleGrey bar: in-sample · Red bar: out-of-sample
Select or focus a bar to inspect a result.
Out-of-sample result
Rule net compared with buy-and-holdRed bar: rule net · Gold line: buy-and-hold
Select or focus a bar to inspect a result.
Rule explorer

Inspect one cell in full.

In-sample PF0.75
Out-of-sample PF0.78
Out-of-sample net-7.38
Buy-and-hold+1.96
Trades610
VerdictFails net

The account finished negative after cost, so the cell failed before benchmark comparison was needed.

The one cell that looked positive

Positive is not the same as passed.

Exactly one cell in the whole grid ended out-of-sample with a positive number. The RSI fade on the thirty-minute timeframe, at measured market cost, finished up 0.28. On a naive reading, that becomes the headline: “RSI fade is profitable out-of-sample.” It is exactly the sentence that can sell a course.

Three pre-committed checks dismantle it. First, it is the only positive cell in twenty, which is entirely consistent with statistical noise. Second, it appears only at the cheaper measured cost. At the original 2.5 bps cost it loses money. Third, the same rule and timeframe lost in-sample with a profit factor of 0.90. Even the positive out-of-sample result trailed buy-and-hold, so it fails the claim regardless.

The useful result is not +0.28.The useful result is that a superficially positive backtest can still fail a frozen, whole-account claim.
Why it failed

Two different failure shapes.

Momentum: no edge to begin with

The moving-average crossover and Donchian breakout sit well below a profit factor of one in both samples. They turn over hundreds of round trips and produce the largest losses. There is no measurable edge here for costs to erase.

Fades: the fitted-history trap

Several fades look plausible in-sample and then fall below one out-of-sample. That shape is consistent with a rule fitted to its own history rather than a repeatable edge.

What this does and does not say

A narrow result, stated firmly.

This Assay is about five specific rules, at frozen parameters, on one pair, at two timeframes, over one window, under two cost treatments. It is not a claim that markets are unbeatable, that all systematic trading fails, or that price-based methods can never work. It says nothing about other pairs, parameters, sessions or windows, and it is not a prediction of the future.

What it does say is small and firm: run Keel's own simple price rules the way a real account pays for them, judge them out-of-sample against a bar fixed in advance, and none survived.

Methodology and evidence

The receipts stay attached.

Every figure traces to a frozen, checksummed evidence capsule containing the configuration, per-cell metrics, receipt and conformance record. The configuration was pre-registered before results existed. The dataset and measured costs were frozen independently. The evaluation was run once and reconciled against the sealed metrics.

Conclusion

Our own rules received no special treatment.

They were frozen before testing, evaluated out-of-sample, charged real costs and reported exactly as they performed. Every cell failed. The result is not that trading is impossible. The result is that evidence must outrank intuition, including our own.

This is independent research on simulated paper execution for educational purposes. It is not financial, investment or trading advice, and nothing here is a recommendation to buy, sell or use any instrument, strategy, system or product.

Cost assumptions are operator measurements and assumptions and must be verified against current venue schedules. Real-world results vary with venue, size and conditions. Past or simulated performance does not indicate future results.

Published evidenceAsk Coach Steve