Rapid forward-testing, quarantined from the main tools. The loop: the miner sweeps (budget stated) → strongest hypotheses auto-register here as paper variants (cap 4, fresh clocks) → +100-checkpoint verdicts kill or continue → a survivor (CLV>0, CI excluding 0, n≥500 — goal G1) gets a leak audit, then goes to Sam. Most candidates SHOULD die.
2360 combinations tested to date — expect ~118.0 false positives at 95% among any winners. This run: 297 tested across 330 distinct combinations ever examined. Bar: n>=30, CLV mean>=2.0%, 95% CI lo>0 (tags: CI lo > implied), AND Benjamini-Hochberg (FDR q=0.05) on the CUMULATIVE look count: one-sided p<=2.12e-05 against a family of 2360 cumulative look(s). Generated 2026-07-25T00:30.
No combination cleared the bar — the honest common case.
Surfaces we have but haven't swept (graduate by becoming miner dimensions):
| field size (per model) small vs capacity fields - different market efficiency? |
| distance bands (per model) sprints vs staying trips |
| venue tier (metro/provincial/country) needs a small venue->tier map (curated, cheap) |
| track grade depth (Good 3 vs Heavy 10) wet/dry is swept; the GRADE within it is not - a Heavy 10 is not a Soft 5, and race_profiles carries the number |
| race number within the meeting the time band is swept, the sequence is not - R1 maidens vs a late-card feature differ in field quality, not just clock time |
| M7 rating vs market disagreement races where M7 and the market disagree most - is either right? |
Data we don't have — and what each would unlock:
| Jockey/trainer strike rates (free (forward) / provider $ (history)) Sam's named M7 feature; FormFav has names only. Accumulates FREE going forward from race_profiles (jockey column, collecting since 2026-07-12) - a backfill needs a provider. |
| Live-key Betfair prices + traded volume (Betfair live key application; engineering moderate) UNLOCKS THE MONEY-FLOW QUARANTINE: real volume, sub-minute prices, true last-10s late money (the T-60s lock's physical floor moves). |
| Full placings (2nd-4th) (on hold per Sam) place-market models, exotic structures, better calibration targets. Quotes in hand: PuntingForm AU$59/mo (DECISION_LOG). |
| Sectional times (PF upgrade $) PF Modeller tier; accessible-tier ratings already tested no-edge, raw sectionals untested. |
| Gear changes / stewards reports (no compliant source identified yet) classic qualitative angles (blinkers first time etc.) - would enter as runner TAGS first. |
| NZ form + weather (moderate) NZ races run blind on form; needs FormFav NZ coverage or provider + MetService stations. |
| Barrier trials / jump-outs (no structured source identified) unraced-horse signal, classic stable-money precursor. |
"Suspect leak" is a holding state, not a verdict. Each entry names the test that would exonerate it.
| miner dimension: model x traffic light (withdrawn 2026-07-23, T-20260722-10) flagged: circular by construction - it conditions CLV on price movement, and CLV IS price movement. The light is the runner's T-90 -> near-jump move relative to its field; clv_pct is its entry-price -> SP move. For any model entering before T-90 those windows overlap almost entirely, so 'it firmed' and 'it beat its entry price' are close to the same sentence. Measured (Analyst-of-record 2026-07-22): mean CLV firm-minus-drift is +40 to +52pp at 90-240min entry lead, +24 to +25pp at T-10, and collapses to nil for the models that enter at T-60s. It is +23.7pp in b1-t10fav-v1, which just backs the T-10 favourite - a 'signal' present in a model with no signal is not a signal, it is the instrument reading itself. exoneration: the light may return as a miner dimension if and only if it is mined against a metric measured ENTIRELY OUTSIDE the light's own price window - e.g. CLV from a T-60s entry price, or strike rate versus market-implied probability - so that the conditioning variable and the metric no longer share a price path. Re-running the same CLV sweep on more races is not an exoneration; the sample was never the problem. |
| +moneyflow vs bsp (delta 0.020715, 2026-07-16) flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap. |
| H:mf_drift_only vs bsp (delta 0.017028, 2026-07-16) flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap. |
| H:mf_vol_surge vs bsp (delta 0.016819, 2026-07-16) flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap. |
| +form+moneyflow vs bsp (delta 0.020058, 2026-07-16) flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap. |
| [adopted] Late money / firming into the jump as a selection signal (became cash-t90/saver-t90 + the light filters) — Sam, 2026-07-19 |
| [open] Tipster formatting signals beyond bold (colour, caps) - capture as tags the day they appear — Sam, 2026-07-19 |
| [open] Monetisation split of the shared dashboard: the public funnel currently exposes the FULL Data Room (full.html + incubator) to friends, which is fine for now - but a later paid/free tier could separate the card (free share) from the full analysis surfaces. Would need: data/mobile restructure into public/ vs premium/ dirs, separate serve ports or auth on the funnel. — Sam, 2026-07-19 |
Add ideas to data/frontier_backlog.json — qualitative welcome; the runner-tag system makes them measurable.