← back to test matrix

Test AH-1

Which big AI is best at drawing a stripe our model missed?

Claude — it recovered 78% of the missing stripes after rules, well ahead of Gemini (61%), Nano Banana (21%), and GPT-5.1 (12%). The ranking held after the rule filter.
What it means for SailScan

If we use a big AI as a rescue step for stripes our model misses, Claude is the one. The rules filter matters: Nano Banana looked decent raw (41%) but lost 43 lines to geometry rules (down to 21%) — it was getting credit for wobbly lines that happened to land near a stripe. Claude and Gemini barely moved, meaning their lines were clean to begin with.

78%
Claude recovered
61%
Gemini
21% / 12%
Nano / GPT
The detail
AIRecovered (raw)Recovered (after rules)Junk / photoLines killed by rules
Claude80%78%0.243
Gemini 3.1 Pro62%61%0.421
Nano Banana Pro41%21%0.2243
GPT-5.112%12%0.882

~88 photos where our model missed a stripe. Each AI was asked to draw the missing stripe; we scored how many it recovered, before and after throwing out lines that break our sail rules.

Source: ai-recovery-rulescored 2026-09-06, ~88 photos, rules from production isGeometryViolation + checkStripeCrossing.

raw data (JSON) →