← back to test matrix

Test AH-2

Can a frontier AI draw the stripes as well as our model?

Not as the main scanner — it found 85% of stripes vs our ~96-97%, and its main failure was declining to trace at all (12 of 37 photos under-counted). But when it did commit, its line accuracy was on par with deep20.
What it means for SailScan

A big general AI is not our stripe-finder today. But 'on par when it commits' greenlights the OTHER roles: judge (score/flag our lines), label factory (pre-trace training photos for humans), and rescue on shaky scans. The honest caveat: our models had seen these photos and the AI hadn't — the fair rematch (AH-3) runs on photos nobody trained on, with a prompt that forces it to commit.

85%
frontier stripes found
96–97%
our models
12/37
photos it under-drew
The detail
Who drew the stripesStripes foundJunk lines / photoClean photos
Frontier AI (cold)85.4%0.4921/37
v26 (our model)96.7%0.3029/37
deep20 (our model)95.9%0.4131/37

37 photos. Big caveat: our models trained on these photos; the frontier AI saw them cold — so this was NOT a fair fight (it favors our models). The fair rematch on untrained photos is test AH-3 / SAIL-149.

Verdict: Not the scanner today; per-line accuracy on par with deep20 when it commits; failure mode = declining to trace (12/37 photos under-counted vs label). Greenlights judge/label-factory/rescue roles; fair rematch belongs on photos no model trained on, with a commit-forcing prompt.

Source: frontier-bench 2026-09-06, 37 holdout-excluded photos. Tracer: claude-fable-5 vision agents. NOTE: incumbents trained on these photos.

raw data (JSON) →