# The 9:28 scanner: every piece of data it collects and uses
As of 2026-09-25 (after the prior-session and premarket-high fixes). Code: `universe.py`, `scanner.py`, `monitor.py`. Rescans (9:40, then every 30 min until 15:15) run the same four steps.

## Step 1: Asset list (Alpaca assets: active US stocks, ~13,500)
| Data | Used for |
|---|---|
| Symbol, tradable flag | Keep tradable stocks only |
| Company name | Drop ETFs/funds/trusts/notes, leveraged/inverse products; keep names of ≤ 5 syllables |
| Config include / exclude lists | Currently empty |
**Result:** ~3,400 symbols.

## Step 2: Snapshot prefilter (one Alpaca snapshot per symbol at 9:28)
| Data | Used for |
|---|---|
| Latest trade price (last premarket trade) | Must be **$5–$15** |
| Previous session's volume | Must be **≥ 500,000**; also the RVOL baseline and Stage 1's volume-pace baseline |
| Previous session's high | Saved as the **previous-day high** level |
| Previous session's close | Saved as the **previous close** (for the gap) |
**Result:** ~300 stocks. (Before the 9/25 fix, these three values came from **two sessions back**.)

## Step 3: Daily bars (last ~40 days)
| Data | Used for |
|---|---|
| 20-day high | Score (resistance) and trade-plan target level |
| Daily ATR (14) | Volatility class; minimum stop = 0.25 × daily ATR |

## Step 4: 1-minute premarket bars (last 6 hours) = the score
| Measure | Calculation | Full credit at | Weight |
|---|---|---|---|
| Relative volume (RVOL) | Premarket volume ÷ previous day's volume | 3× | **30** |
| Premarket volume | Shares traded premarket | 100,000 | **20** |
| Gap | Last premarket price vs previous close (negative if down) | +10% | **20** |
| Premarket price strength | Avg close of last third vs first third of premarket (negative if falling) | +5% | **15** |
| Distance to resistance | Room to nearest level above: premarket high → previous-day high → 20-day high | ≥ 8% room or no level above | **15** |
Scaled to 0–100; **top 30** = watchlist. Also saved: premarket high/low, resistance level used, previous close, premarket volume.

## Used after the scan by the bot
| Saved data | Used by |
|---|---|
| Premarket high, previous-day high, 20-day high | Trade plan: breakout levels and targets |
| Previous day's volume | Stage 1 volume pace (≥ 0.4× normal for the time of day) |
| Daily ATR | Stop floor, volatility class |
| Score / rank | Only the order candidates are checked in |

## Not used at all
Bid/ask spread, float / shares outstanding, news, sector, short interest, multi-day average volume (the "average daily volume" filter is **yesterday's volume only**).

## Which measures actually predicted early runners (20-day study)
**Method:** for every qualifying stock-day from 8/27 to 9/24, every scanner measure was recomputed the way the **fixed** scanner works (correct previous day). The outcome was checked against the first 30 minutes. A **runner** = a 5%+ low-to-high run between 9:30 and 10:00. 3,675 stock-days had premarket trading and could be scored; 303 were runners (8.2%). Script and data: `reports/open_window/scanner_feature_study.py`, `scanner_features_*.csv`.

**Runner rate by each measure** (lowest fifth → highest fifth of stocks):

| Measure | Lowest 20% | Highest 20% | Verdict |
|---|---|---|---|
| **Daily ATR %** (how volatile the stock normally is) | 1.1% | **24.6%** | **Strongest.** Not in the score at all today |
| **Gap size (either direction)** | 4.1% | **23.8%** | **Strong.** But the score **penalizes gaps down**, and gap-downs of 1%+ ran 12.7% of the time |
| RVOL | 2.9% | 22.0% | Strong, but **full credit at 3× is almost unreachable**: 80% of stocks read under 0.02×, so the 30-point weight barely separates them |
| Room to resistance | 3.8–4.5% | 20.8% | Works as intended |
| Premarket volume | 2.7% | 16.9% | Moderate |
| Premarket strength | 3.5% (flat) | 15.9% (up) | **Falling premarkets also ran (12.7%)**: size matters more than direction |
| **Current total score** | 3.8% | 18.9% | Works, but weaker than single measures |

**Share of each day's runners caught by the top 30:**

| Ranking | Runners caught |
|---|---|
| Current score (fixed data) | **39%** (117 of 303) |
| The bot's live lists 9/17–9/24 (buggy data) | 32% (34 of 107) |
| Gap size alone | 53% |
| Daily ATR % alone | 53% |
| Gap size + daily ATR % | 56% |
| **Gap size + daily ATR % + RVOL** (equal weights) | **60%** (183 of 303) |

**What this suggests:**
1. **Add daily ATR %** (normal volatility) to the score. It's the single best early-runner signal.
2. **Score the gap and premarket move by size, not direction.** Gap-downs and falling premarkets also produced early runners (often a bounce).
3. **Recalibrate RVOL** against premarket-sized volume (for example full credit at 0.05–0.10× of yesterday's volume, not 3×).

**Caution:** "caught a runner" isn't the same as "made money". Volatile stocks also fall hard, and a gap-down "run" is often a bounce. A new score should go through the entry backtest (every signal taken, same exits) before replacing the current one. These are simple equal-weight rankings on 20 days, so the risk of overfitting is modest but not zero.
