← back to test matrix

Test 0.9

How often would Nate's 'call the big model' rule fire?

On about 7% of sails (24 of 329). And it's our OWN rules, not the model, causing it — the raw model never drops a sail below 3 stripes.
What it means for SailScan

The gate is worth building: ~1 in 14 sails, a manageable backup volume. The bigger finding underneath: our cleanup rules sometimes over-delete real stripes (one sail went 21 → 1) — a rules-tuning lead on its own.

7.3%
sails end <3 stripes
0%
caused by raw model
24
sails affected
The detail

Nate's idea: only call the expensive big model when the pipeline returns fewer than 3 stripes. This sizes how often that would fire.

StageSails ending <3 stripes
Raw model0 of 329 (0%)
After rules + coherence24 of 329 (7.3%)
The 24 sails that end short
PhotoRealRaw foundFinal
3393ed66-2f9…4211
564c75e1-f65…4160
06e7b1a7-e32…342
21a7c795-74a…332
4e94f98a-2af…352
aca6ec8c-447…332
e27c97c2-2df…392
f8a02b2a-654…352
68ccaf678f1d…370
699649075940…332
699c7862f320…330
69a4c2667974…331
69aeb1c43172…452
69b52c09603c…332
69b826db7a6a…332
69babfb8630c…332
69c44d5c8c71…3131
69c6a1d78c71…342
69cea47eed06…340
69e6c64b5e24…332
69f37dc11c24…331
69f8c879ea23…332
6a414d8fc425…332
c721b1b7-5d7…392

Source: 332-photo holdout, frozen 2026-08-25. Real sail = label has ≥3 stripes.

raw data (JSON) →