Assay 004 · Published · 16 July 2026

We tried to kill the most documented trade in the literature.

The famous version died. The boring one lived.

Every strategy Numerroo Assay had published failed after honest costs. This is the first one that did not, but only in its boring, defensive form.

Category-level premium 9 US-listed ETFs 1993 to June 2026 Frozen evidence attached
The result in one line

Defensive timing, not a way to beat the market.

In one line: Every strategy Numerroo Assay has published so far has failed after honest costs: a crypto grid bot, a Lorentzian classifier, and Keel's own shipped price rules. This is the first one that didn't. We took time-series momentum, arguably the most documented trend premium in the academic literature, and tried to kill it under real retail costs across a nine-asset universe from 1993 to 2026. The literature-faithful long/short version failed: not enough signal after cost, and stone dead in the last decade. The plain retail long/flat version survived every pre-registered test. But read the next sentence before you do anything with that: what survived is defensive timing, not a way to beat the market. It kept about three-quarters of a simple portfolio's excess return while taking under a third of its worst loss: a risk-adjusted result, never an outperformance one.

1 / 2variants survived

The famous long/short construction failed. The retail long/flat construction survived as defensive timing, not market-beating alpha.

Why this one is different

The first survivor matters because the instrument was built to falsify.

Assay exists to charge real costs against strategies that look good on paper and report the result whether it flatters anyone or not. Three published teardowns in, the scoreboard was three falsifications and no survivors. That is the honest track record of the instrument, and it is the reason a survivor here means something: the process was not built to find edges, it was built to try to break them.

So we picked the hardest possible thing to break. Time-series momentum (buy what has gone up over the past year, avoid or short what has gone down) is not an obscure signal. It is one of the most replicated findings in asset pricing, documented across decades, asset classes and continents. If honest retail costs could not touch a premium that well-established, that would be worth knowing. If they could, that would be worth knowing too.

What was tested

Two named variants. No sweeps.

Two named variants of the same signal, frozen before any data was touched. No others were run, and nothing was swept.

Q1 · Literature-faithfulLong / short

The classic academic construction: hold a long position in each asset whose trailing twelve-month return is positive, and a short position in each asset whose trailing twelve-month return is negative. This is the version the papers describe.

Q2 · Retail-implementableLong / flat

The same signal, but the short leg is replaced with cash: go long when the trailing return is positive, sit in the three-month Treasury bill otherwise. This is the version an ordinary investor could actually run in a normal brokerage account, with no borrow, no margin and no short-selling.

The signal is computed at each month-end close and executed at the next trading day's close, never on the same bar. Nine equal-notional slots; monthly cadence; one lookback, one cadence, one weighting, fixed in advance. In Q2, the idle slots earn the T-bill and are never redistributed to the winners. The universe expands as each ETF reaches its inception plus twelve months, so no instrument is used before it existed. The earliest member, SPY, begins in 1993; the portfolio return series covers 389 months to 30 June 2026.

The claim under Assay

All three frozen criteria had to hold.

Stated so it can fail. A variant survives only if all three of these hold:

  1. Full sampleNet-of-cost annualized excess-over-cash return is positive with a t-statistic of at least 2 over the full sample.
  2. Recent decadeNet excess is at least zero over the frozen recent decade, 1 July 2016 to 30 June 2026.
  3. RobustnessThe full-sample sign survives removal of any single instrument from the universe.

All three criteria were fixed and published before the engine ran. Equal-weight buy-and-hold of the same nine assets is reported alongside for context: it is not a pass/fail bar, but it is the honest yardstick for what "good" even means here.

How it was measured

One frozen judge.

Prices

FMP dividend-adjusted daily closes: total-return basis, so dividends are included. The adjustment was verified three ways before any signal code ran, including a no-dividend control (GLD) that came back at exactly zero adjustment as expected.

Cash

The FRED three-month T-bill series (TB3MS), used for the excess-return calculation and for Q2's flat leg.

Retail scale

US$5,000 per slot, roughly US$45,000 deployed, at an IBKR-class broker. Costs are a frozen, source-dated table: half the published bid/ask spread plus the IBKR Pro fixed commission, charged on every unit of traded notional. Per-instrument one-way costs run about 2.5 to 5 basis points. Where a real issuer-published spread was captured it was used; where it was not, a deliberately high ceiling was frozen instead, so the cost assumptions bias against the strategy surviving.

Publication rule

Datasets, costs and criteria were sealed and checksummed before the result was computed. The engine ran once. The numbers below are what it returned, read as-is.

The result

One signal. Two very different outcomes.

Q1, the famous long/short version, fails. Over the full sample it earned +2.31% per year in net excess return, but with a t-statistic of only 1.49, short of the required 2, so that return is not statistically distinguishable from luck. Worse, over the recent decade it earned -0.11% per year: essentially nothing, with the short book actively bleeding. It passed the robustness check (positive with any single asset removed) but failed two of three criteria. The trend-skepticism that has dogged trend-following since the 2009 low shows up exactly where you would expect it: in the shorts.

Q2, the plain retail long/flat version, survives. Full sample: +3.56% per year in net excess, t-statistic 3.37. Recent decade: +2.82% per year, still clearly positive. Robustness: positive with any single instrument removed, the weakest case being +2.57% per year with SPY dropped. All three frozen criteria pass. Doubling every trading cost as a stress test barely moves it (+3.53% per year), confirming that at monthly cadence, transaction cost is a footnote, not the story.

The single change that flips the verdict is the short leg. Replace shorting with cash and the strategy goes from failing to surviving. The longs' timing carried; the shorts destroyed it. Read across the two variants, the short book cost roughly 2.9% per year over the recent decade.

Result window

Showing the full 389-month sample.

Selected Assay 004 results
VariantAnnualized excesst-statisticMax drawdownWinning monthsCriterion
Q1 long/shortLiterature-faithful+2.31%1.49−40.5%59.6%Fails full sample
Q2 long/flatRetail-implementable+3.56%3.37−12.5%60.4%Passes full sample
Net excess return
Annualized result by variantSwitch the result window above. Zero is the recent-decade pass line.
Select or focus a bar to inspect a result.
Full-sample signal
t-statistic against the frozen thresholdOnly the full-sample criterion requires t ≥ 2.
Select or focus a bar to inspect a result.
Variant explorer

Inspect one construction in full.

Annualized excess+3.56%
t-statistic3.37
Max drawdown−12.5%
Winning months60.4%
Full samplePass
Recent decadePass
RobustnessPass
Overall verdictSurvives

Q2 cleared all three frozen criteria, but buy-and-hold still earned more raw excess return.

What “survived” means, and what it does not

Defensive timing, not market-beating alpha.

This is the part to slow down on, because "survived a pre-registered test" is very easy to mis-read as "beats the market." It does not.

Over the same 389 months, simply buying and holding all nine assets in equal weight earned more raw excess return (about +4.80% per year) than the surviving strategy's +3.56%. If the only thing you cared about was total return, buy-and-hold won.

What the strategy bought you was not more return. It was a far smoother ride. Buy-and-hold's worst peak-to-trough loss over the period was about -45%. Q2's was about -12.5%. So the surviving strategy kept roughly three-quarters of buy-and-hold's excess return while suffering under a third of its worst drawdown. On a risk-adjusted basis it is genuinely better; on a raw-return basis it is not. The only honest claim here is the risk-adjusted one. Anyone who summarizes this as "trend-following beats the market" has misread the evidence.

Q2 retained74%

of buy-and-hold excess return

Q2 suffered28%

of buy-and-hold drawdown

Q2 long/flat+3.56%

annualized excess · −12.5% max drawdown

Buy-and-hold+4.80%

annualized excess · −44.9% max drawdown

Three caveats you must read first, not last

The survivor is conditional.

  1. Defensive, not enhancing.As above: this is drawdown control, not a return engine. It underperformed simple buy-and-hold on raw return. Its value, if any, is that it lost far less in the bad years.
  2. The survival is sensitive to what idle cash earns.Q2's flat leg is parked in Treasury bills. If that cash instead earned nothing (a 0% cash assumption), the full-sample result drops to +2.10% per year with a t-statistic of 1.97, a hair under the pass bar. The strategy is robust to doubled trading costs but fragile to the yield on idle cash, which is broker- and balance-dependent and not something every retail investor actually captures. This is a genuine near-miss and we report it prominently, not in a footnote.
  3. The universe is survivorship-tilted and the spreads are modern.The nine ETFs are all liquid and alive today; a universe chosen in 1993 might have included funds that later closed. And today's tight spreads are applied across the whole history, which understates what trading actually cost in the 1990s and 2000s. The doubled-cost stress test is our attempt to bound that second point; the first is disclosed and unmodelled.
Assumption stress

Cash yield, not trading cost, is the fragile point.

Select or focus a bar to inspect a sensitivity.
What this does and does not say

One narrow result, kept inside its boundary.

It says

One specific, pre-registered, honestly-costed construction of a documented premium cleared a bar we set before we looked, as a risk-adjusted, defensive strategy, in its retail long/flat form only, on a modern liquid universe. It says the famous long/short form did not clear that bar, and died in the last decade.

It does not say

This beats buying and holding. It does not say it will survive out-of-sample from today, on a different universe, at a different scale, or with realistic idle-cash yields. It does not say anyone should trade it. Under Numerroo's own rules, a surviving backtest is not a trading decision: it is, at most, a candidate for honest forward paper-tracking, which is a separate decision that has not been made.

Methodology and evidence

The receipt is attached.

Everything behind the numbers above is frozen and checksummed. The evidence capsule is:

  • Receipt: assay-004-tsmom-multiasset-1993-2026-v1
  • Evidence-engine commit: f9b0694
  • Manifest SHA-256: 1e9979758b54c646a1c0917f46b66e9e8be42538dd9c96d6c5a621a89b9f3541
  • Artifacts: result.json (the read of record), matrix.json (the frozen month-end price and cash inputs), b6_costs.py (the source-dated cost table).

The result is not merely hash-consistent, it is reproducible: re-running the frozen engine at commit f9b0694 against the frozen matrix.json regenerates result.json identically, field for field. The 0%-cash and doubled-cost sensitivities quoted above are context figures derived from the same frozen inputs by the same engine, published with their own derivation receipt so every number on this page traces to sealed evidence.

What happens next

A survivor is not a trading decision.

Per the pre-registration, a surviving variant does not become a live or paper strategy automatically. Q2 is parked as an explicit decision point, not a next step. If it is ever tracked forward, that will be pre-registered on its own terms (rules frozen before data, as everything here was) and reported here whether it holds up or not.

This is research, not financial advice, and not a recommendation to buy, sell, hold or trade anything. Past results, especially backtested ones, do not predict future returns. Numerroo Assay sells no signals, bots, strategies or courses, and takes no broker, bot or course affiliate revenue. The evidence is frozen and public so anyone can check the arithmetic against the claims.

Published evidenceAsk Coach Steve