The buy-the-dip setup that works — but won't fit in a portfolio
Grade A — pre-registered before testing, held-out out-of-sample split, reproduced to the decimal.
Status — PROVED (signal-level) — live as a watchlist (/pullback) · score v1
The claim
The classic swing-trade setup: find a strong stock in a clear uptrend, wait for it to pull back to support (the 50-day moving average), and buy when it reclaims — bounces back above and closes strong. "Buy the dip, but only in leaders." It's one of the most-taught setups in trading education, and unlike a lot of them, it has an intuitive logic: you're buying a proven winner at a temporary discount, not catching a falling knife.
This is the one strategy in our founding set we tested the right way from the start — the hypothesis, the universe, the window, and the exact pass/fail bar were all written down before the first backtest ran. So it earns our highest evidence grade, and it's the standard the rest now follow.
The verdict, in one line
The signal genuinely works out-of-sample — and that turns out not to be enough. Each individual setup carries a real, positive edge that survived a strict held-out test. But the setup fires far more often than any real account has room for, and when we're forced to choose which ones to take, the ones we skip quietly beat the ones we keep. A great signal you can't fully hold.
How we tested it
- Universe: S&P 500 (503 names), 2010–2026.
- Signal: leaders above a rising 200-day average, near their 1-year high, that pull back to tag the 50-day and then post a bullish reclaim bar.
- The test, pre-registered: measure the average outcome per trade (in "R" — multiples of the risk taken) on a training era (2011–2019) and a held-out test era (2020–2026). Written in advance: pass if expectancy ≥ +0.30R in both eras with no sign flip; kill if the test era is ≤ 0 or the sign flips. Same exit engine as our other strategies, so only the entry differs.
- We re-ran this study on 2026-07-03 and it reproduced to the decimal.
What we found
Stage 1 — the signal passed cleanly:
| Era | Trades | Avg outcome per trade | Bar |
|---|---|---|---|
| TRAIN 2011–2019 | 8,984 | +0.862R | +0.30R |
| TEST 2020–2026 (held out) | 5,121 | +0.552R | +0.30R |
Both eras clear the bar comfortably, with no sign flip. The test-era number is lower than training (honest decay of −0.31R), but +0.55R per trade on 5,000 held-out trades is a real, durable signal. That is a genuinely rare result — most setups that look good in-sample collapse out-of-sample.
Stage 2 — then reality's ceiling. We ran the passing signal as an actual portfolio, and it badly lagged a plain index hold in the test era (+59% vs the market's +148%). The cause isn't signal quality — it's capacity. The strategy fires ~5,000 test-era signals into an account with room for maybe 50 positions at a time. When you can only take a fraction, which fraction matters enormously — and here's the sting: the trades we were forced to skip averaged +0.63R, while the ones we took averaged only +0.32R. Picking first-come, we systematically left the better trades on the table.
The fix that didn't work
The obvious answer: train a model to pick the best signals. So we did — a proper machine-learning study with a pre-registered bar (improve selection by ≥ +0.10R), a purged walk-forward, and one shot at the test set.
It failed, and it failed in the most instructive way: the model's "best-ranked" trades (+0.37R) underperformed its "worst-ranked" ones (+0.64R). It learned the relationship backwards. The lesson we now hold as a house view: the information needed to pick winners is not in the price chart. No amount of clever modelling on price-and-volume features recovered it.
The one thing that did help, a little, was adding genuinely new information — dark-pool / short-volume order-flow data lifted selection by a modest but real +0.145R out-of-sample. That's why the live /pullback page carries a Score column built from order flow, not from price. The honest principle: the path forward is new data, not new models.
Honest caveats
- 1.Capacity is the whole story. As a hand-picked watchlist, the +0.55R signal edge is real and usable. As a mechanical portfolio, it doesn't beat the index — don't run it that way.
- 2.Adverse selection is real. Even choosing carefully, you may skip the better trades. The order-flow score helps at the margin; it does not solve it.
- 3.One test era, survivorship-limited. 503 of today's index names; the absolute numbers are flattered by survivors, as with any free-data equity study.
How the score breaks down
| Dimension | Score | Why |
|---|---|---|
| Edge | 16 / 25 | A clear, OOS-validated signal edge (+0.55R held out) — but it does not convert to portfolio outperformance over the index, because of capacity. Scored as a real signal, docked for the deployment ceiling. |
| Robustness | 17 / 25 | Passed a strict held-out test with no sign flip; gated and ungated arms agree. Docked for a single test era and the stage-2 capacity fragility. |
| Practicality | 15 / 25 | Ships and works as a manual watchlist (its live home). Not a set-and-forget portfolio — it needs hand-selection, and even then adverse selection bites. |
| Evidence | 22 / 25 | Pre-registered with a written pass/kill bar, a genuinely held-out test set, honest decay reported, and reproduced to the decimal. The best-evidenced strategy we have. |
| Total | 70 / 100 | Qualified edge — a real signal, best-in-class evidence, held back by a capacity ceiling we've been unable to model our way around. |
Evidence Grade A: the rare one we tested properly from the first run. A lower total than some weaker-evidenced strategies would score is exactly the point — the grade and the score answer different questions.
What it means for you
Treat pullback as a watchlist, not an autopilot. The individual setup is one of the few we've found with a real out-of-sample edge, which makes it a strong source of ideas to review by hand. Use the order-flow score to break ties — but know that picking is genuinely hard here, and that the strategy's biggest limitation isn't the signal, it's how many good ones it hands you at once.
References
- Jegadeesh & Titman (1993) on momentum in leaders; the pullback is a mean-reversion entry within an uptrend, a widely-taught discretionary setup with no single canonical paper.
- The four-dimension scoring rubric behind the number above: The Tapelab Score.
Scored under score v1, the fixed public rubric summarised on the research index. Report last updated 2026-07-03; published from the research repository's canonical write-up, unedited. Nothing here is investment advice.