Defensive timing, not a way to beat the market.
In one line: Every strategy Numerroo Assay has published so far has failed after honest costs: a crypto grid bot, a Lorentzian classifier, and Keel's own shipped price rules. This is the first one that didn't. We took time-series momentum, arguably the most documented trend premium in the academic literature, and tried to kill it under real retail costs across a nine-asset universe from 1993 to 2026. The literature-faithful long/short version failed: not enough signal after cost, and stone dead in the last decade. The plain retail long/flat version survived every pre-registered test. But read the next sentence before you do anything with that: what survived is defensive timing, not a way to beat the market. It kept about three-quarters of a simple portfolio's excess return while taking under a third of its worst loss: a risk-adjusted result, never an outperformance one.
The first survivor matters because the instrument was built to falsify.
Assay exists to charge real costs against strategies that look good on paper and report the result whether it flatters anyone or not. Three published teardowns in, the scoreboard was three falsifications and no survivors. That is the honest track record of the instrument, and it is the reason a survivor here means something: the process was not built to find edges, it was built to try to break them.
So we picked the hardest possible thing to break. Time-series momentum (buy what has gone up over the past year, avoid or short what has gone down) is not an obscure signal. It is one of the most replicated findings in asset pricing, documented across decades, asset classes and continents. If honest retail costs could not touch a premium that well-established, that would be worth knowing. If they could, that would be worth knowing too.
Two named variants. No sweeps.
Two named variants of the same signal, frozen before any data was touched. No others were run, and nothing was swept.
The classic academic construction: hold a long position in each asset whose trailing twelve-month return is positive, and a short position in each asset whose trailing twelve-month return is negative. This is the version the papers describe.
The same signal, but the short leg is replaced with cash: go long when the trailing return is positive, sit in the three-month Treasury bill otherwise. This is the version an ordinary investor could actually run in a normal brokerage account, with no borrow, no margin and no short-selling.
The signal is computed at each month-end close and executed at the next trading day's close, never on the same bar. Nine equal-notional slots; monthly cadence; one lookback, one cadence, one weighting, fixed in advance. In Q2, the idle slots earn the T-bill and are never redistributed to the winners. The universe expands as each ETF reaches its inception plus twelve months, so no instrument is used before it existed. The earliest member, SPY, begins in 1993; the portfolio return series covers 389 months to 30 June 2026.
All three frozen criteria had to hold.
Stated so it can fail. A variant survives only if all three of these hold:
- Full sampleNet-of-cost annualized excess-over-cash return is positive with a t-statistic of at least 2 over the full sample.
- Recent decadeNet excess is at least zero over the frozen recent decade, 1 July 2016 to 30 June 2026.
- RobustnessThe full-sample sign survives removal of any single instrument from the universe.
All three criteria were fixed and published before the engine ran. Equal-weight buy-and-hold of the same nine assets is reported alongside for context: it is not a pass/fail bar, but it is the honest yardstick for what "good" even means here.
One frozen judge.
FMP dividend-adjusted daily closes: total-return basis, so dividends are included. The adjustment was verified three ways before any signal code ran, including a no-dividend control (GLD) that came back at exactly zero adjustment as expected.
The FRED three-month T-bill series (TB3MS), used for the excess-return calculation and for Q2's flat leg.
US$5,000 per slot, roughly US$45,000 deployed, at an IBKR-class broker. Costs are a frozen, source-dated table: half the published bid/ask spread plus the IBKR Pro fixed commission, charged on every unit of traded notional. Per-instrument one-way costs run about 2.5 to 5 basis points. Where a real issuer-published spread was captured it was used; where it was not, a deliberately high ceiling was frozen instead, so the cost assumptions bias against the strategy surviving.
Datasets, costs and criteria were sealed and checksummed before the result was computed. The engine ran once. The numbers below are what it returned, read as-is.
One signal. Two very different outcomes.
Q1, the famous long/short version, fails. Over the full sample it earned +2.31% per year in net excess return, but with a t-statistic of only 1.49, short of the required 2, so that return is not statistically distinguishable from luck. Worse, over the recent decade it earned -0.11% per year: essentially nothing, with the short book actively bleeding. It passed the robustness check (positive with any single asset removed) but failed two of three criteria. The trend-skepticism that has dogged trend-following since the 2009 low shows up exactly where you would expect it: in the shorts.
Q2, the plain retail long/flat version, survives. Full sample: +3.56% per year in net excess, t-statistic 3.37. Recent decade: +2.82% per year, still clearly positive. Robustness: positive with any single instrument removed, the weakest case being +2.57% per year with SPY dropped. All three frozen criteria pass. Doubling every trading cost as a stress test barely moves it (+3.53% per year), confirming that at monthly cadence, transaction cost is a footnote, not the story.
The single change that flips the verdict is the short leg. Replace shorting with cash and the strategy goes from failing to surviving. The longs' timing carried; the shorts destroyed it. Read across the two variants, the short book cost roughly 2.9% per year over the recent decade.
Showing the full 389-month sample.
| Variant | Annualized excess | t-statistic | Max drawdown | Winning months | Criterion |
|---|---|---|---|---|---|
| Q1 long/shortLiterature-faithful | +2.31% | 1.49 | −40.5% | 59.6% | Fails full sample |
| Q2 long/flatRetail-implementable | +3.56% | 3.37 | −12.5% | 60.4% | Passes full sample |
Inspect one construction in full.
Q2 cleared all three frozen criteria, but buy-and-hold still earned more raw excess return.
Defensive timing, not market-beating alpha.
This is the part to slow down on, because "survived a pre-registered test" is very easy to mis-read as "beats the market." It does not.
Over the same 389 months, simply buying and holding all nine assets in equal weight earned more raw excess return (about +4.80% per year) than the surviving strategy's +3.56%. If the only thing you cared about was total return, buy-and-hold won.
What the strategy bought you was not more return. It was a far smoother ride. Buy-and-hold's worst peak-to-trough loss over the period was about -45%. Q2's was about -12.5%. So the surviving strategy kept roughly three-quarters of buy-and-hold's excess return while suffering under a third of its worst drawdown. On a risk-adjusted basis it is genuinely better; on a raw-return basis it is not. The only honest claim here is the risk-adjusted one. Anyone who summarizes this as "trend-following beats the market" has misread the evidence.
of buy-and-hold excess return
of buy-and-hold drawdown
annualized excess · −12.5% max drawdown
annualized excess · −44.9% max drawdown
The survivor is conditional.
- Defensive, not enhancing.As above: this is drawdown control, not a return engine. It underperformed simple buy-and-hold on raw return. Its value, if any, is that it lost far less in the bad years.
- The survival is sensitive to what idle cash earns.Q2's flat leg is parked in Treasury bills. If that cash instead earned nothing (a 0% cash assumption), the full-sample result drops to +2.10% per year with a t-statistic of 1.97, a hair under the pass bar. The strategy is robust to doubled trading costs but fragile to the yield on idle cash, which is broker- and balance-dependent and not something every retail investor actually captures. This is a genuine near-miss and we report it prominently, not in a footnote.
- The universe is survivorship-tilted and the spreads are modern.The nine ETFs are all liquid and alive today; a universe chosen in 1993 might have included funds that later closed. And today's tight spreads are applied across the whole history, which understates what trading actually cost in the 1990s and 2000s. The doubled-cost stress test is our attempt to bound that second point; the first is disclosed and unmodelled.
Cash yield, not trading cost, is the fragile point.
One narrow result, kept inside its boundary.
It says
One specific, pre-registered, honestly-costed construction of a documented premium cleared a bar we set before we looked, as a risk-adjusted, defensive strategy, in its retail long/flat form only, on a modern liquid universe. It says the famous long/short form did not clear that bar, and died in the last decade.
It does not say
This beats buying and holding. It does not say it will survive out-of-sample from today, on a different universe, at a different scale, or with realistic idle-cash yields. It does not say anyone should trade it. Under Numerroo's own rules, a surviving backtest is not a trading decision: it is, at most, a candidate for honest forward paper-tracking, which is a separate decision that has not been made.
The receipt is attached.
Everything behind the numbers above is frozen and checksummed. The evidence capsule is:
- Receipt:
assay-004-tsmom-multiasset-1993-2026-v1 - Evidence-engine commit:
f9b0694 - Manifest SHA-256:
1e9979758b54c646a1c0917f46b66e9e8be42538dd9c96d6c5a621a89b9f3541 - Artifacts:
result.json(the read of record),matrix.json(the frozen month-end price and cash inputs),b6_costs.py(the source-dated cost table).
The result is not merely hash-consistent, it is reproducible: re-running the frozen engine at commit f9b0694 against the frozen matrix.json regenerates result.json identically, field for field. The 0%-cash and doubled-cost sensitivities quoted above are context figures derived from the same frozen inputs by the same engine, published with their own derivation receipt so every number on this page traces to sealed evidence.
A survivor is not a trading decision.
Per the pre-registration, a surviving variant does not become a live or paper strategy automatically. Q2 is parked as an explicit decision point, not a next step. If it is ever tracked forward, that will be pre-registered on its own terms (rules frozen before data, as everything here was) and reported here whether it holds up or not.
This is research, not financial advice, and not a recommendation to buy, sell, hold or trade anything. Past results, especially backtested ones, do not predict future returns. Numerroo Assay sells no signals, bots, strategies or courses, and takes no broker, bot or course affiliate revenue. The evidence is frozen and public so anyone can check the arithmetic against the claims.