RacingAI · M8

Model Incubator

← back to the card

Rapid forward-testing, quarantined from the main tools. The loop: the miner sweeps (budget stated) → strongest hypotheses auto-register here as paper variants (cap 4, fresh clocks) → +100-checkpoint verdicts kill or continue → a survivor (CLV>0, CI excluding 0, n≥500 — goal G1) gets a leak audit, then goes to Sam. Most candidates SHOULD die.

Active candidates

auto-tipster2-firm-v1 ACTIVE
mirrors tipster2-v1 · filter [light=firm] · born 2026-07-19
Scored 63 (+0 pending) · wins 17 · CLV -3.6% ±3.3 · P&L -8.5u · next checkpoint n=100

Latest sweep

2360 combinations tested to date — expect ~118.0 false positives at 95% among any winners. This run: 297 tested across 330 distinct combinations ever examined. Bar: n>=30, CLV mean>=2.0%, 95% CI lo>0 (tags: CI lo > implied), AND Benjamini-Hochberg (FDR q=0.05) on the CUMULATIVE look count: one-sided p<=2.12e-05 against a family of 2360 cumulative look(s). Generated 2026-07-25T00:30.

No combination cleared the bar — the honest common case.

Frontier — what to look at next

Surfaces we have but haven't swept (graduate by becoming miner dimensions):

field size (per model)
small vs capacity fields - different market efficiency?
distance bands (per model)
sprints vs staying trips
venue tier (metro/provincial/country)
needs a small venue->tier map (curated, cheap)
track grade depth (Good 3 vs Heavy 10)
wet/dry is swept; the GRADE within it is not - a Heavy 10 is not a Soft 5, and race_profiles carries the number
race number within the meeting
the time band is swept, the sequence is not - R1 maidens vs a late-card feature differ in field quality, not just clock time
M7 rating vs market disagreement
races where M7 and the market disagree most - is either right?

Data we don't have — and what each would unlock:

Jockey/trainer strike rates (free (forward) / provider $ (history))
Sam's named M7 feature; FormFav has names only. Accumulates FREE going forward from race_profiles (jockey column, collecting since 2026-07-12) - a backfill needs a provider.
Live-key Betfair prices + traded volume (Betfair live key application; engineering moderate)
UNLOCKS THE MONEY-FLOW QUARANTINE: real volume, sub-minute prices, true last-10s late money (the T-60s lock's physical floor moves).
Full placings (2nd-4th) (on hold per Sam)
place-market models, exotic structures, better calibration targets. Quotes in hand: PuntingForm AU$59/mo (DECISION_LOG).
Sectional times (PF upgrade $)
PF Modeller tier; accessible-tier ratings already tested no-edge, raw sectionals untested.
Gear changes / stewards reports (no compliant source identified yet)
classic qualitative angles (blinkers first time etc.) - would enter as runner TAGS first.
NZ form + weather (moderate)
NZ races run blind on form; needs FormFav NZ coverage or provider + MetService stations.
Barrier trials / jump-outs (no structured source identified)
unraced-horse signal, classic stable-money precursor.

Quarantine — flagged, not buried

"Suspect leak" is a holding state, not a verdict. Each entry names the test that would exonerate it.

miner dimension: model x traffic light (withdrawn 2026-07-23, T-20260722-10)
flagged: circular by construction - it conditions CLV on price movement, and CLV IS price movement. The light is the runner's T-90 -> near-jump move relative to its field; clv_pct is its entry-price -> SP move. For any model entering before T-90 those windows overlap almost entirely, so 'it firmed' and 'it beat its entry price' are close to the same sentence. Measured (Analyst-of-record 2026-07-22): mean CLV firm-minus-drift is +40 to +52pp at 90-240min entry lead, +24 to +25pp at T-10, and collapses to nil for the models that enter at T-60s. It is +23.7pp in b1-t10fav-v1, which just backs the T-10 favourite - a 'signal' present in a model with no signal is not a signal, it is the instrument reading itself.
exoneration: the light may return as a miner dimension if and only if it is mined against a metric measured ENTIRELY OUTSIDE the light's own price window - e.g. CLV from a T-60s entry price, or strike rate versus market-implied probability - so that the conditioning variable and the metric no longer share a price path. Re-running the same CLV sweep on more races is not an exoneration; the sample was never the problem.
+moneyflow vs bsp (delta 0.020715, 2026-07-16)
flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured
exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap.
H:mf_drift_only vs bsp (delta 0.017028, 2026-07-16)
flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured
exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap.
H:mf_vol_surge vs bsp (delta 0.016819, 2026-07-16)
flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured
exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap.
+form+moneyflow vs bsp (delta 0.020058, 2026-07-16)
flagged: beat the BSP close - impossible for tradeable pre-race info on the data as captured
exoneration: re-test on LIVE-KEY takeable prices (not WAPs) - if the effect survives on real executable prices, it was never a leak. Blocked on the live-key data gap.

Ideas backlog

[adopted] Late money / firming into the jump as a selection signal (became cash-t90/saver-t90 + the light filters) — Sam, 2026-07-19
[open] Tipster formatting signals beyond bold (colour, caps) - capture as tags the day they appear — Sam, 2026-07-19
[open] Monetisation split of the shared dashboard: the public funnel currently exposes the FULL Data Room (full.html + incubator) to friends, which is fine for now - but a later paid/free tier could separate the card (free share) from the full analysis surfaces. Would need: data/mobile restructure into public/ vs premium/ dirs, separate serve ports or auth on the funnel. — Sam, 2026-07-19

Add ideas to data/frontier_backlog.json — qualitative welcome; the runner-tag system makes them measurable.

generated 19:58 · paper only · promotion beyond paper is always Sam's call