This sets the ceiling for a junk-filter approach (test 1.3): at best it can win back about half of today's misses, not all of them. That's still real accuracy on the table, worth building — but the other half needs a fundamentally better proposer, not just a smarter filter. Greenlit test 1.3 with that ceiling in mind.
| Stripe position | Missed by our model | Sitting in SAM3's pool |
|---|---|---|
| Top | 14 | 8 |
| Middle | 14 | 6 |
| Bottom | 30 | 15 |
| Total | 58 | 29 |
SAM3-solo draws far more candidate lines than we keep (an "over-propose" pool of 2,838 lines across 329 photos). This test checks: of the stripes our production model misses entirely, how many of them are hiding somewhere in that oversized pool, just waiting on a filter to pick them out?
Source: 332-photo holdout, frozen 2026-08-25; SAM3-solo pool from runs/hopeval/D-sam3-solo.