A big general AI is not our stripe-finder today. But 'on par when it commits' greenlights the OTHER roles: judge (score/flag our lines), label factory (pre-trace training photos for humans), and rescue on shaky scans. The honest caveat: our models had seen these photos and the AI hadn't — the fair rematch (AH-3) runs on photos nobody trained on, with a prompt that forces it to commit.
| Who drew the stripes | Stripes found | Junk lines / photo | Clean photos |
|---|---|---|---|
| Frontier AI (cold) | 85.4% | 0.49 | 21/37 |
| v26 (our model) | 96.7% | 0.30 | 29/37 |
| deep20 (our model) | 95.9% | 0.41 | 31/37 |
37 photos. Big caveat: our models trained on these photos; the frontier AI saw them cold — so this was NOT a fair fight (it favors our models). The fair rematch on untrained photos is test AH-3 / SAIL-149.
Verdict: Not the scanner today; per-line accuracy on par with deep20 when it commits; failure mode = declining to trace (12/37 photos under-counted vs label). Greenlights judge/label-factory/rescue roles; fair rematch belongs on photos no model trained on, with a commit-forcing prompt.
Source: frontier-bench 2026-09-06, 37 holdout-excluded photos. Tracer: claude-fable-5 vision agents. NOTE: incumbents trained on these photos.