Scan US large-cap equities for smooth uptrends (high trailing return paired with shallow drawdown) and track which names persist across runs. Use when the user wants to find what's working in the market, scan for momentum, discover the next NVDA / LITE / MU-style breakout before headlines, spot leading sectors or themes (AI infra, semis, defense, lithium, etc.), surface persistent winners across runs, or compare current leaders to a prior run. Also covers re-runs and parameter tweaks ("run it...
Installs into .claude/skills of the current project.
Are you the author of Momentum Scan?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/mthli-momentum-scan)
---
name: momentum-scan
description: Scan US large-cap equities for smooth uptrends (high trailing return paired with shallow drawdown) and track which names persist across runs. Use when the user wants to find what's working in the market, scan for momentum, discover the next NVDA / LITE / MU-style breakout before headlines, spot leading sectors or themes (AI infra, semis, defense, lithium, etc.), surface persistent winners across runs, or compare current leaders to a prior run. Also covers re-runs and parameter tweaks ("run it again", "anything new showing up", "3 month window", "include small caps"). Do NOT use for single-ticker price or fundamentals lookups, ETF holdings, chart generation, value-investing screens, or generic explanations of momentum investing; those need other tools or plain answers.
---
# momentum-scan
Find US equities in **smooth uptrends** (high trailing return with shallow drawdown) and surface which names are durable leaders vs single-week pops. The value over a one-shot screener is **persistence tracking**: the script logs each US market day (America/New_York) once to `state/history.csv` (re-running the same day refreshes that day's snapshot rather than appending), so each subsequent run can compute streak, rank changes, dropouts, and new entrants.
By default each run also surfaces two **entry-timing layers** on top of the momentum filter: a **pullback entry signal** (MA20 distance + RSI(14) β π’ buy zone / π΅ deep pullback / π‘ in trend / π stretched / π΄ overextended) that flags whether each pick is buyable now vs already extended, and an **ATR-based stop loss** (2.5Γ ATR by default) for per-position risk sizing. The pullback signal answers "is this buyable right now?", the canonical complement to momentum's "what's running?" question: momentum names tend to arrive already 30-50% above MA20, a state where mean-reversion pullbacks often give back a meaningful slice of the gain before the trend resumes.
A **vol-collapse filter** (`--vol-collapse-ratio`, default 0.2) also runs after the score-based ranking (before persistence enrichment) to catch the canonical signature of an acquisition target pinned at the announced cash offer price: the announcement-day gap inflates the window return while the post-event flat tape shrinks the max drawdown, and together they yield an outlier Score you can't trade as momentum. The same signature catches reverse-merger / SPAC lock-ins. Excluded names get a dedicated output section with their pre/post-event annualized vol so you can sanity-check the trigger. Deep documentation (banner/JSON schema, excluded-ticker lifecycle, window-position sensitivity, false positives) lives in `references/vol-collapse.md`; read it when an exclusion fires or a name goes missing.
**Dependencies** (auto-fetched by `uv run --with`): Python β₯ 3.10, `yfinance>=1.3,<2`, `pandas>=2` (the script uses `format="ISO8601"`, added in pandas 2.0), `numpy>=1.24,<3`. No persistent venv needed.
`<SKILL_DIR>` below is the directory containing this `SKILL.md`. Substitute the absolute path when running.
## Run
```bash
# Standard run: 3mo window, top 30
uv run --with 'yfinance>=1.3,<2' --with 'pandas>=2' --with 'numpy>=1.24,<3' \
python <SKILL_DIR>/scripts/scan.py
# Longer window for smoother, slower-moving leaders
... python <SKILL_DIR>/scripts/scan.py --window-months 6
# Inspect the run history (no new scan)
... python <SKILL_DIR>/scripts/scan.py --show-history
# Machine-readable JSON output
... python <SKILL_DIR>/scripts/scan.py --format json
# Strict trend filter: suppress top-N when SPY is below a rising 200DMA
... python <SKILL_DIR>/scripts/scan.py --regime-gate strict
# Vol-targeted sizing: cohort 60d vol β leverage; per-name Weight% column
... python <SKILL_DIR>/scripts/scan.py --target-vol-pct 15
# Override ATR stop multiplier (default 2.5; pass 0 to disable)
... python <SKILL_DIR>/scripts/scan.py --atr-stop-mult 3.0
# Skip pullback entry indicator (no MA20% / RSI / Sig columns)
... python <SKILL_DIR>/scripts/scan.py --no-pullback
# Skip sector tagging (faster first run, no Sector column or breakdown line)
... python <SKILL_DIR>/scripts/scan.py --no-sectors
# Disable the vol-collapse acquisition-target filter (keep buyouts in the table; see parameter table for range)
... python <SKILL_DIR>/scripts/scan.py --vol-collapse-ratio 0
# Skip the benchmark refresh a saving run does by default (~10s of price fetching)
... python <SKILL_DIR>/scripts/scan.py --no-benchmark
# Rebuild the board-vs-index curves on their own (scan.py already does this after
# each history save; run it by hand after --no-benchmark, or to change --board-n / --top-n).
... python <SKILL_DIR>/scripts/compute_benchmark.py # --board-n 10 --top-n 30 --refresh-prices
# Render history.csv + sectors.json (+ benchmark.json) into a self-contained HTML
# dashboard at state/history.html (benchmark panel, rank bump chart, sector stacks,
# heatmap, per-ticker table). Stdlib-only; plain python, no uv --with needed.
python <SKILL_DIR>/scripts/render_history_html.py # --top-n 30 --days 60 --out <path>
# Outcome backtest: replay history.csv's episodes (enter on listing, sell on dropout),
# stratified by entry attributes (see "Backtested outcomes" section). Re-run quarterly.
... python <SKILL_DIR>/scripts/backtest_outcomes.py
# Realistic execution variant: both fills at the NEXT session's open
... python <SKILL_DIR>/scripts/backtest_outcomes.py --fills next-open
```
## Parameters
| Flag | Default | Notes |
|---|---|---|
| `--window-months` | 3 | Lookback for return + max drawdown. Shorter = earlier signals, more noise. Bump to 6 for smoother, slower-moving leaders; drop to 1β2 to capture an in-progress sector rotation the 3mo window is still too slow to show (see Known limitations on short-window behavior). The minimum-observations guard is `min(60, trading_days - 3)` and applies end-to-end (both the close-extraction step and `score_tickers` share it), so it scales with the window: 1mo needs ~18 sessions, 2mo ~39, and β₯3mo keeps the historical 60-session floor unchanged. |
| `--top-n` | 30 | Display cutoff. History logs **every** name that passed the filter each day (below-cutoff rows preserve near-miss context), but the persistence stats (Streak / FirstSeen / RankΞ / π) only count appearances at rank β€ N, i.e. runs where the name was displayed. A name that hovered at #40 for weeks still debuts as π with streak 1. Changing `--top-n` between runs re-interprets which historical appearances count as visible. |
| `--min-return-pct` | 30 | Filter floor on trailing return over the window. |
| `--max-dd-pct` | 20 | Filter ceiling on max drawdown (absolute value). |
| `--min-market-cap` | 5e9 | Universe market-cap floor. Lower = include small-cap rockets but more noise. |
| `--min-volume` | 1e6 | Universe avg-3mo-volume floor (liquidity filter). |
| `--universe-count` | (all matches) | Universe size pulled from Yahoo's screener. Default unset = pull every match the screener reports (currently ~1000 US large caps at default mcap/volume floors). The screener returns at most **250 rows per request** (Yahoo's hard cap; `yf.screen` raises `ValueError` above that), so the script paginates the universe in 250-row pages with `offset`; at default filters that's ~5 paginated requests, taking a few extra seconds, but only on cache refresh (every 7 days). Pass an explicit positive integer to cap the universe at the top-N largest by market cap (e.g. `250` for a one-request refresh, `500` for the previous default size); argparse rejects 0 / negative values. If you raise `--universe-count` above the number of tickers already in `state/universe.txt`, the script force-refreshes the cache even within TTL; otherwise you'd keep getting the smaller cached pool with no warning. In the other direction, the script slices a cache *larger* than the requested count to the first `count` names without a refresh (the file orders names by descending market cap, so the slice keeps the top-count largest). Older yfinance versions (without `offset` support) fall back to a single 250-row page. |
| `--refresh-universe` | (auto, 7d TTL) | Force-refresh universe (ignore cache). |
| `--no-refresh-universe` | β | Use cached universe even if past TTL (offline / testing). |
| `--show-history` | β | Dump history summary, no new scan. Applies the `--top-n` visibility cutoff to its streak / frequency / climber stats, same as a live run. |
| `--clear-history` | β | Wipe `state/history.csv`. |
| `--no-save` | β | Run but don't append to history (useful for one-off exploration). |
| `--save-stale` | β | Override the non-trading-day guard. By default the script skips `append_history` when today's ET date is a weekend or NYSE-observed holiday so streak counts don't inflate from duplicate-data days. Pre-market runs on a real trading day still save. |
| `--allow-same-day` | β | Keep existing rows for today's ET date instead of overwriting them (debugging / forcing multiple snapshots). |
| `--prune-non-trading-days` | β | One-shot cleanup: drop history rows whose ET-date `run_date` is not an NYSE trading day. Use after upgrading from a pre-guard version, or after intentional `--save-stale` runs. Runs no scan. |
| `--format` | markdown | `markdown` or `json`. |
| `--verbose` | β | Restore the diagnostic columns (AnnVol%, RankΞ, FirstSeen, FromHigh%, MA20%, RSI) to the top-N table. The default slim table (2026-07-31 redesign) keeps the decision columns β Trend sparkline, return, drawdown, Score, Streak, Sig, Stop β since Sig already encodes MA20/RSI and the Trend trajectory supersedes RankΞ/FirstSeen. JSON always carries every field. |
| `--regime-gate` | warn | Market trend filter. `off` skips it entirely (and the longer data fetch). `warn` shows a SPY/breadth banner + a RISK-OFF caveat; top-N still printed. `strict` suppresses the top-N when RISK-OFF (history is still saved so streaks survive). RISK-ON means SPY > 200DMA *and* the 200DMA slope over the last 20 trading days is above a small `-0.05%` dead band (so a near-flat MA doesn't flip on single-bar noise). |
| `--target-vol-pct` | (off) | Portfolio vol target in % (e.g. `15` for 15% annualized). When set: computes the equal-weight cohort's 60-day realized vol, surfaces `suggested leverage = target / cohort_vol` (clipped to `[0.25, 1.0]`, deleverage-only per Daniel-Moskowitz 2016), and adds a `Weight%` column using equal-risk-contribution Γ leverage. The weights sum to `leverage Γ 100`, so a 0.6Γ leverage means you hold 60% notional and 40% cash. Off = no Weight% column. |
| `--atr-stop-mult` | 2.5 | ATR-based stop multiplier. Computes 14-day ATR for each top-N pick, adds a `Stop` column showing `last_close - mult Γ ATR` as both price and % from spot. Names with `Streak β₯ --persistent-min-streak` also get a `TrailStop` line in the Persistent leaders section, anchored to the peak since `FirstSeen`. Typical multipliers: `2.0` tight (frequent stop-outs, lower per-trade loss), `2.5` standard (default), `3.0` loose (rarer stop-outs, larger per-trade loss). Pass `0` or a negative value to disable the Stop column. |
| `--no-pullback` | β | Disable the pullback entry indicator. Default behavior computes MA20 distance and RSI(14, Wilder) for each top-N pick and shows three columns: `MA20%` (price relative to its 20-day average), `RSI` (14-day Wilder RSI), and `Sig` (π’/π΅/π‘/π /π΄ classification). Evaluation order is π’ β π΅ β π΄ β π β π‘ (first match wins). π’ = MA20% in [-3, +3] *and* RSI in [40, 55] (classic Trend Pullback buy zone). π΅ = MA20% β€ -3 *and* RSI < 40 (Connors-style deep pullback in a still-intact uptrend: strong risk/reward *if* the trend holds, but harder to confirm than π’ since it can also be the leading edge of a broken trend). π΄ = MA20% > 25 *or* RSI > 80 (overextended; chasing here tends to give back a meaningful slice on the first pullback). π = MA20% in (15, 25] or RSI in (70, 80] (stretched, wait). π‘ = everything else (in trend, neutral). Pair with the momentum filter to surface buyable-right-now names rather than already-extended runners. |
| `--persistent-min-streak` | 3 | Streak threshold used by both the **Persistent leaders** section and the ATR `TrailStop`. Default `3` matches the historical display threshold. Bump to `4` if you only want streaks that have survived multiple periods of noise, the "real signal" cutoff from the interpretation guide. |
| `--no-benchmark` | β | Skip rebuilding `state/benchmark.json` after the history save. Default: after saving history, the scan refetches closes for every name that ever made the board (plus SPY / QQQ, one batched request, ~10s) and rewrites the board-vs-index curves the dashboard's benchmark panel draws: the board curve holds each day's top 10 (`compute_benchmark.py --board-n`), and the price block covers **this run's** `--top-n`. Runs after the report is printed, so it never delays the table, and no failure in it can end the run. Skipped when history isn't saved (`--no-save`, or a non-trading day without `--save-stale`), and skipped when an existing `benchmark.json` was computed for a different `--top-n` β the tracked file keeps following the board the dashboard draws, and `compute_benchmark.py --top-n` is how you switch it. |
| `--no-sectors` | β | Disable sector tagging. Default: fetches sector/industry from yfinance for the top-N picks (cached in `state/sectors.json`, 30-day TTL), shows a `Sector` column in the table and a `**Sectors**` breakdown line in the header. First-run cost is ~1β2s per uncached pick (parallelized at 10 workers); subsequent runs hit the cache. Pass `--no-sectors` to skip it. |
| `--vol-collapse-ratio` | 0.2 | Acquisition-target / lock-in filter. Computes annualized realized vol for the **first** and **second** halves of the scoring window and excludes names where `vol_second / vol_first < ratio` (with a `vol_first β₯ 5%` annualized floor; names already at very low vol have an unstable ratio). Default 0.2 catches the canonical signature (announcement gap + post-event price-pin β 2nd-half vol collapses to single digits) while leaving normal consolidating-after-a-rally names alone. **Raise** to `0.3` (hard cap `1.0`) to catch more lock-in patterns at the cost of more false positives on calm-drifting earnings-pop names. **Lower** to `0.15` to require a more dramatic collapse (fewer exclusions, may leak buyouts through). Pass `0` or negative to disable. See `references/vol-collapse.md` for the banner/JSON schema, excluded-ticker lifecycle, and sensitivity/false-positive analysis. |
## Output shape
A markdown table of the top N, plus three discovery sections. Sample (truncated):
```
# Momentum scan β 2026-05-11 16:07 UTC
**Params**: window=3mo, min_return=30.0%, max_dd=20.0%, mcap>5e+09
**Universe**: ~1000 tickers Β· **Passed filter**: ~50 (vol-collapse: 0 excluded) Β· **Prior runs**: 1
**Data**: daily bars through 2026-05-11
**Regime**: SPY 612.4 vs 200DMA 558.2 (+9.7%) Β· 50DMA > 200DMA Β· 200DMA slope (20d): +0.18% Β· Breadth: 68% > 200DMA β **RISK-ON**
**Vol target**: cohort 60d vol 24.8% β suggested leverage **0.60x** (target 15%, raw 0.60x, clip 0.25β1.00x)
**Sectors**: Technology 11 Β· Energy 5 Β· Healthcare 4 Β· Communication Services 3 Β· Industrials 2 Β· Other 5
**Sig**: π’2 π΅0 π‘23 π 5 π΄0
## Top 10
| # | Ticker | Sector | Trend | 3m% | MaxDD% | Score | Streak | Sig | Stop | Weight% |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | **NOK** | Tech | βββ | +92.4 | -8.0 | 11.6 | 3 | π | $5.42 (-5.5%) | 2.6 |
| 2 | **MRVL** | Tech | π | +109.4 | -10.8 | 10.1 | 1 | π‘ | $82.10 (-6.8%) | 1.5 |
| 3 | **DELL** | Tech | ββββββββββ | +107.8 | -10.8 | 10.0 | 1 | π | $128.40 (-7.1%)| 1.5 |
| 4 | **AMD** | Tech | βββββββββ β | +115.0 | -11.6 | 9.9 | 2 | π΄ | $231.10 (-7.4%)| 1.4 |
| 5 | **CIEN** | Tech | ββββββ | +103.5 | -16.8 | 6.2 | 6 | π‘ | $108.30 (-9.2%)| 1.2 |
...
_Trend: rank trajectory, last β€10 runs Β· β = #1 Β· β = #30 or worse Β· rising = climbing Β· full diagnostic columns: --verbose_
## Dropouts since last run (6)
- **ASX** (was #2, 3m=+126.6%)
- **TTE** (was #3, 3m=+45.8%)
...
## New entrants (6)
_entry quality by entry-day distribution days (the validated edge): π’ clean β€1 Β· βͺ mixed 2-3 Β· π loaded 4+. Suffix +surge/+quiet = entry-day volume β₯1.5Γ/<0.8Γ, a weak secondary signal_
- βͺ **MRVL** at #2 (3m +109.4%, MaxDD -10.8%) Β· 3 dist, vol 2.1Γ
- π **DELL** at #3 (3m +107.8%, MaxDD -10.8%, re-entry, was #24) Β· 5 dist, vol 0.6Γ
...
## Persistent leaders (streak β₯ 3 runs)
- `ββββββ` **CIEN**: streak 6, first seen 2026-04-21, now #5 Β· trail stop $108.50 (-9.3% from spot, peak $120.30)
```
The **Sig** strip under the banner is the cohort-level buyability read (see interpretation point 5): a π΄-heavy strip means the whole cohort is extended; a shift toward π’/π΅ usually means the correction already happened. The table is the **slim** default (2026-07-31 redesign); `--verbose` restores the AnnVol% / RankΞ / FirstSeen / FromHigh% / MA20% / RSI diagnostic columns, and JSON always carries every field.
(The section lists **episode starts** (`Streak = 1`), covering both first-ever debuts (π in the table) and re-entries after a dropout; re-entries carry a `re-entry, was #N` note. Entry-quality tags tier each entrant **by its entry-day distribution-day count** (`dist_days_25d`), the half of the signal the backtest validated (β€1-dist-day entrants roughly doubled tenure and top-10 reach; **Backtested outcomes** #2): π’ clean β€ 1 dist Β· βͺ mixed 2-3 Β· π loaded β₯ 4. The entry-day volume character (`vol_ratio_20d`) is only a label suffix (`+surge` β₯ 1.5Γ, `+quiet` < 0.8Γ) because its original calibration was convention-inflated (**Backtested outcomes** #3; tiers were volume-primary before 2026-07-31). Read the tag as a priority hint for which entrants deserve attention. Tags appear whenever `dist_days_25d` is available; the volume suffix also needs `vol_ratio_20d`. JSON output carries the tier on each episode-start pick as `entry_quality: {emoji, label}` (labels like `clean+surge`, `mixed`, `loaded+quiet`).)
(The `trail stop ...` suffix only appears when `--atr-stop-mult` is set *and* `Streak β₯ --persistent-min-streak`, which controls both the Persistent leaders threshold and the trail-stop attach threshold. Names below it skip the suffix.)
The **Trend** column in the table and the leading `` `βββββ` `` prefix in Persistent leaders are the same encoding: the name's **leaderboard-rank trajectory** over its last β€10 appearances (from `score_rank`, current run appended; unlike Streak it includes below-cutoff days, which clamp to the floor block). Orientation is inverted (taller block = better rank, so a rising line means *climbing*), and heights use a fixed 1..`top_n` scale, so blocks are comparable across names (#1 is always tallest). With fewer than two data points the Trend cell shows π. Read-only display nicety; affects no scoring or saved state.
Numbers above are illustrative; real `Stop%` spans roughly -5% (sedate) to -22% (high-vol breakouts). The sample doesn't include a π΅ row because deep pullbacks in still-strong momentum names are uncommon; they tend to follow sharp short-term sell-offs in otherwise-trending leaders.
Column meanings (columns marked *verbose* print only with `--verbose`; JSON always carries them):
- **Trend**: rank-trajectory sparkline (see above). Supersedes the numeric `RankΞ`/`FirstSeen` pair in the slim table: `βββββββββ` reads as "just surged into the leaderboard", `ββββββββββ` as "entrenched leader", `βββ βββββββ` as "former leader bleeding out".
- **AnnVol%** (*verbose*): annualized realized volatility over the scoring window (`--window-months`). Useful both as a per-name sanity check (a +100% return at 60% vol is a coin flip held the right way; the same return at 25% vol is a real trend) and as the input to the per-name `Weight%` allocation when `--target-vol-pct` is set.
- **Score**: return Γ· |max drawdown|. Higher = more return per unit of pain.
- **Streak**: consecutive prior runs this ticker was in the *displayed* top N (1 = first appearance). History also stores below-cutoff rows (every name that passed the filter), but those don't count toward Streak / FirstSeen / RankΞ / π; the stats only reference ranks you saw.
- **RankΞ** (*verbose*): `(score_rank at latest prior top-N appearance) β (current score_rank)`. Positive β = rising; negative β = slipping; π = no prior top-N appearance in the entire history. Note: the "latest prior appearance" can predate the previous run: a ticker that fell out for a few runs and is now back shows the delta against its last-seen rank, not against the previous run (where it was absent). The `FirstSeen` column and the **New entrants** / **Dropouts** sections cover the "in-and-out" view; RankΞ stays focused on "how has this name moved since we last saw it". **Important**: the script computes the delta on `score_rank` (the pre-vol-collapse score-based rank; see Output shape), not on the display rank `#` in the leftmost column. This means a vol-collapse exclusion of a top pick won't produce false `+1 β` for every name below it. In rare cases when vol-collapse status flips between adjacent runs for the same name (e.g., a name re-passing the filter after a few days excluded), the displayed `prev_rank` and the new display `#` may not match RankΞ arithmetic; that's intentional (RankΞ measures real score movement, while `prev_rank` shows what the user saw last time).
- **FirstSeen** (*verbose*): earliest date this ticker appeared in a past run's displayed top-N. Still shown in the Persistent leaders section either way.
- **Weight%**: only present with `--target-vol-pct`. Equal-risk-contribution weight (β 1/ann_vol) scaled by the suggested leverage. The column sums to `leverage Γ 100` (so e.g. 60 means 60% notional, 40% cash). Treat it as a sizing *starting point* rather than a target portfolio; see Known limitations on the correlation simplification.
- **Sector**: only present when sector tagging is enabled (default; disable with `--no-sectors`). Abbreviated GICS-ish sector from yfinance. The `**Sectors**` breakdown line above the table gives full names and counts.
- **MA20%, RSI** (*verbose*) **and Sig**: only present when the pullback indicator is enabled (default; disable with `--no-pullback`); the slim table keeps `Sig`, which encodes both numerics. `MA20%` is `(last_close / MA20 - 1) Γ 100`: positive means above the 20-day average, negative means below. `RSI` is the 14-day Wilder RSI (canonical, EWMA with Ξ±=1/14). `Sig` is the buy-zone classifier (π’ buy / π΅ deep pullback / π‘ watch / π stretched / π΄ overextended); see the `--no-pullback` row in the parameter table for thresholds. Reading rule: π’ and π΅ are *candidates worth investigating today*; π /π΄ are quality momentum but you're late: set a price alert at MA20 and wait.
- **Stop**: present by default (disable by passing `--atr-stop-mult 0`). Format: `$price (-%)`. The stop price = `last_close - mult Γ ATR(14)`. The % shows how far below the current price that sits. For names with `Streak β₯ --persistent-min-streak` (default 3), the Persistent leaders section also surfaces a `TrailStop` anchored to the peak since `FirstSeen` (locks in profits as the trend matures).
The script computes the three discovery sections (dropouts / new entrants / persistent leaders) against the most recent prior run. When the vol-collapse filter removes one or more names, an **Excluded by vol-collapse filter** section appears *between the Regime banner and the Top-N table*. This placement is deliberate: the exclusions are warnings about names that *look* like momentum but aren't, and they print even when `--regime-gate strict` suppresses the rest of the output, so the user always sees them. The section lists each ticker with its `1st-half% β 2nd-half%` annualized vol and the resulting ratio, so you can sanity-check the trigger and recognize the underlying situation (most often a cash buyout pending shareholder vote).
JSON output mirrors the markdown structure: the top-level envelope has a `picks` array (kept entries, with `rank` = display position 1..N and `score_rank` = pre-filter score ordering) and an `excluded_vol_collapse` array (excluded entries, with `rank: null`, `pre_filter_rank` = where they sat in the score ordering, and the `vol_first_pct` / `vol_second_pct` / `vol_ratio` triple). `rank_delta` on kept picks comes from `score_rank` (not display rank), so removing a top pick by vol-collapse doesn't make every name below it show a false +1 β. The `--show-history` view applies the same score_rank-aware delta when computing biggest climbers / droppers across run pairs (falls back to display rank for old history rows without the column).
### Vol-collapse filter: further reading
`references/vol-collapse.md` documents the banner format variants, the excluded-entry JSON schema, the short-window caveat, and the full lifecycle of an excluded ticker (why it appears in Dropouts exactly once, where to look for it afterward, what survives when it returns). Read it when an exclusion fires or the user asks about a missing name.
## How to interpret (Claude's job after running)
The script gives you data; the user wants signal. Add a short interpretation pass: apply judgment rather than reciting the principles below.
Relay the script's markdown output **in full β every row of the top-N table and every section**. Don't truncate or collapse rows to save space: the bottom of the table carries its own signal (names sliding out, names hovering at the cutoff), and the user chose the slim table precisely so the whole thing fits.
Write the interpretation for a reader with **no finance background**, in the conversation's language. Translate each term the moment you use it β "momentum" means "stocks that have been rising steadily"; "Streak 48" means "on the leaderboard 48 scan-days in a row"; "Stop $96.92 (-14.2%)" means "if you owned this, the exit point that caps the loss at 14%". Lead with the story the glyphs tell (which theme is crowding the leaderboard, who's surging in, who's bleeding out), and close with what to do about it β usually "nothing today; these are names to research, not buys".
Before the numbered points, check the `**Data**` line: it names the session the numbers reflect. A `β οΈ Stale data` warning means Yahoo hadn't published the last session's daily bars and the intraday rebuild failed too, so the whole board, every close and every stop is a session old. Open the summary with that, in plain words ("these numbers are from Thursday's close, not Friday's"). A `rebuilt from 30m intraday bars` note is the workaround succeeding (closes within ~0.1% of the official ones). Only that session's volume is unknown, so entrants carry no `+surge` / `+quiet` suffix and the session can't count as a distribution day; no need to mention it unless an entrant's tag matters to the answer.
1. **Read the Regime banner first.** SPY above a *rising* 200DMA with breadth above ~60% is where long-momentum has shown the cleanest risk/reward in the historical record. RISK-OFF banners (SPY below 200DMA, the 200DMA itself rolling over, or breadth collapsing while SPY still holds up) flip the read: treat the names below as *who's holding up* in a weak tape, not *what to buy*. Say up front that the filter doesn't defend against the post-bear momentum crashes (2009 Q2, 2020 Q2, early 2023); those hit right after the gate turns back on, when investors sell prior leaders to fund the rotation into the bombed-out cohort. The filter helps with bear-market downside, not with the regime-flip itself.
2. **Sector clusters beat individual names.** Momentum arrives as a theme (AI infra, semis, defense, lithium, etc.). Group the top 10β15 by sector and call out the cluster; that's what the user can research, hedge, or fade. New entrants joining an existing cluster confirm the theme; isolated newcomers in unrelated sectors are more likely noise.
3. **Streak β₯ 4 and top-5 dropouts are the real signals.** Long streaks have survived multiple periods of market noise; these are the durable trends the rank score alone can't surface. The Persistent leaders section uses a `β₯ 3` threshold by default to surface emerging stickiness early; bump `--persistent-min-streak 4` to filter to only the high-conviction names. A name leaving the top 5 tends to mark a broken trend (max drawdown blew through the filter) and is often the leading edge of a regime shift.
4. **Vol-target the cohort, not individual names.** When `--target-vol-pct` is set, lead with the cohort vol and suggested leverage; that's the antidote to the post-bear momentum crash the trend filter misses. A cohort vol drifting up while leverage drops from 1.0x to 0.4x is the vol-target system *working*: it deleverages into the storm. The per-name Weight% column is useful but secondary; emphasize the leverage number.
5. **Lead with the pullback Sig column when calling out candidate names.** When pullback is enabled (default), the `Sig` column is the single most decision-relevant cell per row; use it as a *research-priority filter*, not a buy signal:
- **π’**: top priority for research today; setup geometry tends to pair with the tighter end of the cohort's Stop% range (RSI cooling and realized-vol cooling arrive together, which compresses ATR).
- **π΅**: also actionable but harder to confirm: a deep pullback in a still-trending name can be the best risk/reward setup *if* the long-term trend is intact, or the leading edge of a broken trend. Cross-check `Streak β₯ 3` and sector stability before recommending.
- **π **: research the underlying, set a price alert at MA20, don't initiate today.
- **π΄**: skip; add to watchlist. Stop% runs wider on high-vol breakouts, but a π΄ triggered by RSI alone on a low-vol name (e.g., a defensive sector pop where ATR stayed small) can still have a tight stop; read Stop% per row rather than assuming it from the Sig color.
A cohort dominated by π΄ (typical of late-stage momentum runs after several uninterrupted up-weeks) is itself a signal: the cohort as a whole is overextended and even healthy names will get sold in a market wobble. When the cohort flips toward π’/π΅/π‘, the likely read is that the broader correction already happened and survivors have rebased.
6. **When ATR stops are on, use Stop% as the buy-time risk number.** With `--atr-stop-mult 2.5`, Stop% runs `-5 to -10%` for sedate names (UNH, large-cap utilities) and `-15 to -22%` for high-vol breakouts (ARM, INTC during a +150% run). The wide range reflects that ATR is name-specific, by design. That % is the *per-trade* max loss if you enter at current price and the stop holds. For dollar-position sizing, `shares = risk_per_trade / (mult Γ ATR)`; that sizing is independent of (and complementary to) the vol-target Weight% column, which sizes by *portfolio* risk. The TrailStop on persistent-leader names (streak β₯ `--persistent-min-streak`) is the more important number for already-running positions: it answers "where would I cut this without giving back the gain". Stops moving up week-over-week is the trend confirming itself.
7. **When sectors are on, lead with the cluster, not individual names.** The `**Sectors**` line is the most user-actionable single piece of info: Tech 12 of 23 means the cohort is concentrated and the vol-target Weight% column is understating true portfolio risk (correlated names). If the top sectors form one cluster (e.g. AI infra: Tech + parts of Comm Svc), say so. Diversifying across sectors at the same total leverage tends to beat chasing the highest-Score name.
8. **Never recommend specific buys.** Frame results as "names worth investigating", not "you should buy". Flag that momentum strategies carry multi-year underperformance risk: 2023 was a textbook momentum crash where the 2022 leaders (energy) lost to a different cohort (mega-cap tech) for the entire year.
## State files
Full schemas, write semantics and sizes are in `references/state-files.md`. Read it before editing a state file by hand, when a column's meaning matters to an answer, or when the `**Data**` line reports a rebuilt or stale session.
- `state/history.csv`: one snapshot per US market day Γ every name that passed the filter. Tracked in git; streak, dropouts and the backtest all read it. `data_asof` records the session the run's bars reached, which equals `run_id` on a healthy run.
- `state/benchmark.json`: the board-vs-SPY/QQQ curves plus the price block behind the dashboard's hover cards, rebuilt after every history save. Tracked.
- `state/history.html`: the dashboard, regenerated from the two files above. Gitignored.
- `state/sectors.json`, `state/universe.txt`: caches with 30-day and 7-day TTLs.
`--clear-history` wipes only `history.csv`, with no confirmation prompt. Tests live next to the scripts:
```bash
uv run --with 'yfinance>=1.3,<2' --with 'pandas>=2' --with 'numpy>=1.24,<3' \
--with 'pytest' pytest scripts/
```
## Backtested outcomes (2026-05-14 β 2026-07-30 sample)
`scripts/backtest_outcomes.py` replays `state/history.csv` under the skill's canonical convention (enter at the close of the first top-30 day, sell at the close of the dropout-observation day) and stratifies by entry attributes. 228 episodes over 50 run-days, one regime; re-run quarterly. Findings, strongest first; full evidence, magnitudes, conventions, and caveats in `references/backtest-findings.md`:
1. **Sell the dropout; don't "give it a few days"**: dropped names keep falling for ~2 weeks, former top-10 names worst.
2. **Entry-day distribution days are the durable entry edge**: clean (β€1 dist) entrants roughly double tenure and top-10 reach vs loaded (4+).
3. **The volume-surge tag was convention-inflated**: the honest convention keeps the surge>quiet ordering but shrinks the gap ~12Γ (5.9pt β 0.5pt).
4. **Entry Score buys persistence, not entry-point return**: top-tercile scores triple tenure but give back the most at the exit; size by score, enter on pullbacks (`Sig`).
5. **Closed episodes are the losers by construction**: the carry is skew (dropouts cut fast while open winners ride); judge the system on both halves.
### Should the trailing stop be a rule?
Not yet, and the dashboard is careful not to imply otherwise. The trailing stop drawn in the hover cards (highest close since entry β mult Γ ATR) answers a sizing question β how much of an open gain is still exposed β and finding #1 above is the only exit this skill has evidence for. Every price- and volume-based exit rule tested against that baseline lost to it, and the climax-type rules were actively harmful; an ATR trail belongs to the same family, so the prior is against it.
It has not, however, been tested *as an ATR trail*, and now it can be: `benchmark.json`'s `px` block is exactly the per-day close and ATR such a replay needs. That is the real reason the number is on the page β persisting the inputs is what turns a hunch into a question with an answer. Add an `--exit-mode atr-trail` to `backtest_outcomes.py` and settle it on the next quarterly re-run.
One fidelity caveat to carry into that test: the dropout exit is decided on a close (the scan runs after the bell), while a real trailing stop is an intraday touch. A trail backtested on closes is the optimistic version of the rule, and the gap widens with volatility.
## Cadence
Cadence-agnostic by design. The script keeps at most one snapshot per US market day (America/New_York), so streak counts **consecutive prior scan-days** containing this ticker; running twice on the same ET day refreshes that day's entry. Aligning to ET date instead of UTC matches what the underlying data represents (US market sessions) and stays stable across DST transitions. The script skips the history write on weekends and NYSE-observed market holidays so streak doesn't inflate from duplicate-data days; results still print, with nothing appended. Pre-market runs on a real trading day **do** save: today counts as a trading day from the streak's perspective regardless of run time. Override the guard with `--save-stale`. To clean up snapshots saved on non-trading days after the fact (pre-guard, or after `--save-stale`), run `--prune-non-trading-days`. The `FirstSeen` dates tell you the natural granularity:
- Daily runs β streak unit is days. Finest granularity, but it adds limited signal over weekly.
- Weekly runs β streak unit is weeks. Recommended sweet spot: captures trend formation 4Γ faster than monthly while smoothing daily noise.
- Monthly runs β streak unit is months. Smoothest, slowest signal; matches the cadence of the original backtest in this conversation.
For automatic recurring runs, use a local scheduler (macOS `launchd` LaunchAgent, or `cron`) pointed at `scripts/scan.py`. The `schedule` skill runs *remote* agents in Anthropic-managed sandboxes that can't see this local `state/` directory, so it doesn't fit this use case.
## Known limitations
- **Survivorship bias**: the universe is current US large caps; delisted names (Lehman, SVB, etc.) are absent. Backtested CAGR runs 1β2% optimistic vs a true point-in-time universe.
- **Pre-cost**: no transaction costs, slippage, or taxes modeled. Real execution shaves another ~0.5β1% CAGR.
- **mcap floor at $5B**: the default floor excludes small-cap moonshots; bump `--min-market-cap 1e9` if the user wants to see them.
- **3mo window is noisier**: fresher-breakout signals come with more single-week pops; bump `--window-months 6` for smoother trends if needed.
- **Short windows (1β2mo) answer a different question**: a 1mo window surfaces *what the last four weeks of flow is rotating into*, often a different cohort than the 3mo durable leaders. Expect a stretched, π /π΄-heavy table, Β±100 `RankΞ` swings, and the *return* floor (not the drawdown ceiling) becoming the binding filter; to widen a thin 1mo table, lower `--min-return-pct`, not `--max-dd-pct`. Full mechanics (drawdown-window geometry, the vol-collapse floor, the historical zero-picks bug) in `references/short-windows.md`.
- **Yahoo data quirks**: rare missing bars, occasional late dividend adjustments. If a single name looks wrong, sanity-check it via the `yfinance` skill's `fast_info` mode. Since 2026-09-02 the just-closed session's daily bar often isn't out by the evening scan; the scan rebuilds it from intraday bars and says so on the `**Data**` line (mechanics, accuracy and the `data_asof` column in `references/state-files.md`).
- **Trend-filter limits**: `--regime-gate` reduces *bear-market* downside but can't catch the regime-flip momentum crash itself (2009/2020/2023); the slope check defuses single-bar whipsaws but two consecutive months of choppy 200DMA crossings can still flip the verdict twice. The `Breadth` figure is the % of the current ~1000-large-cap universe above its own 200DMA, a tech-tilted read of internals rather than a true market-wide A/D line.
- **Vol target ignores correlations**: `--target-vol-pct` uses the equal-weight cohort's *realized* portfolio vol (which does capture correlations through the historical basket return) but the per-name `Weight%` allocation uses `1/ann_vol_i` *as if names were independent*. When the top-N is dominated by one cluster (e.g., AI infra), true portfolio risk is higher than the weights imply. Counter by lowering `--target-vol-pct` (treat 15% as if it were 10% when the cohort is concentrated), or by capping per-name weight by hand. The 60-day SMA lookback also lags: vol regime shifts show up with a delay; an EWMA would react faster but adds a tunable.
- **Vol-target lookback is fixed at 60 trading days, independent of `--window-months`**: the per-name `AnnVol%` in the table is annualized over the *scoring* window (3mo by default, 6mo if you bump it), but the cohort vol driving the leverage calculation stays fixed at the last 60 trading days. This is intentional: the literature (Daniel-Moskowitz, Barroso-Santa-Clara) converges on ~60d as the regime estimator that best predicts the next-period momentum crash. If you bump `--window-months 6`, expect the per-name vol numbers to drift while the cohort vol / leverage stay anchored to the shorter view.
- **ATR stop is per-trade, not portfolio**: `--atr-stop-mult` sizes one position's loss-on-stop; it doesn't account for correlated drawdowns across the top-N. If 20 names share a sector cluster and the cluster sells off, all stops can fire together for a portfolio-wide loss far above any single Stop%. Combine with `--target-vol-pct` (portfolio sizing) for both axes of protection. The 14-day SMA ATR also lags: fast regime changes (gap-down opens, news shocks) blow through computed stops; treat Stop% as a *risk budget*, not a guarantee.
- **Pullback Sig is a buyability filter, not a directional forecast**: π’/π΅ means "if you wanted to buy this name, this is a reasonable entry"; it does *not* mean the name will go up next. π΄ doesn't mean "sell"; it means "don't initiate a new long here". On strong-trending names, the Sig can sit in π΄ for weeks without the price correcting (RSI stays > 70 in a true uptrend); the Trend Pullback edge is *probability* of a better entry showing up, not certainty. RSI thresholds (40-55 buy zone, > 80 overext) are conventional literature defaults that behave well on liquid US large-caps. When you lower `--min-market-cap`, the failure modes cut both ways: small-caps and biotech run RSI > 90 in single-stock squeezes without mean-reverting (false π΄), and illiquid small-caps can sit at MA20 with RSI 50 for weeks without going anywhere (false π’). The indicator detects setup *geometry*, not price follow-through. Lookback is 20 trading days for MA and 14 for RSI, fixed rather than exposed as flags.
- **Sectors are Yahoo, not GICS**: `state/sectors.json` mirrors yfinance's `Ticker.info.sector` / `.industry`, which approximates GICS but differs in spots (e.g., Yahoo uses "Technology"/"Financial Services" where GICS uses "Information Technology"/"Financials"). The abbreviation map handles the common variants; new labels fall through as the first 10 chars of the raw string. Some tickers (ADRs, recent IPOs) return empty sector strings and show as `β` in the table; they don't contribute to the `**Sectors**` breakdown count.
- **Universe pagination has a hard stop at `SCREENER_MAX_PAGES` (20 pages = ~5000 tickers)**: this only matters if Yahoo's response stops including the `total` field (schema drift), in which case the script falls back to per-page heuristics (short page / zero new tickers) to detect end-of-results. The 20-page cap is the absolute backstop; if it triggers you'll see `refresh_universe: hit SCREENER_MAX_PAGES=20 backstop` on stderr and the universe caps at ~5000, well above any realistic large-cap match count, so triggering it signals something wrong upstream.
- **Universe size affects historical rank comparability**: if you raise `--universe-count` (or the underlying universe grows from market-cap drift) between runs, ranks recorded in `state/history.csv` from before the change aren't comparable to ranks after. A larger universe means more names can pass the filter, which can demote previously-high-ranked names because new entrants joined the pool, with no change in the names themselves. `Streak` and `FirstSeen` survive (a ticker is still "in the top N" or not), but read `RankΞ` across a universe-size change with that caveat. If you want clean before/after comparison, run `--clear-history` after changing universe size.
- **Vol-collapse filter sensitivity depends on where the gap day lands in the window**: a gap in the window's *second* half leaks through (late detection rather than a silent miss: the gap drifts into the first half within weeks and gets caught), and 6mo windows can leak names the 3mo scan filters. The lock-in fingerprint when you suspect a leak: `FromHigh% = 0.0` **and** `MaxDD% > -3%`; confirm via the `yfinance` skill's `sec_filings --type PREM14A,DEFM14A`. Both failure directions, the MASI case study, and mitigations in `references/vol-collapse.md`.