← back to test matrix

Test 0.1

Does adding our sail rules help or hurt?

Yes — the rules turn a junky model into a clean one: junk lines drop from ~10% to ~2%, giving up only about 4 real stripes in 100.
What it means for SailScan

The raw model finds nearly every stripe but draws a lot of junk. Our rules layer throws most of that junk out, so what a customer sees is trustworthy. We give up a little recall for far fewer wrong lines — the right trade for a tool people rely on.

90.2%
junk-free BEFORE rules
97.8%
junk-free AFTER rules
~4%
stripes given up
The detail
StageRecall
real stripes found
Precision
drawn lines that are real
Avg line error
Model alone99.8%90.2%0.0078
+ our rules95.8%97.8%0.0114
+ coherence95.8%97.8%0.0114

Recall = how many real human-drawn stripes the stage still finds. Precision = how many of its lines are real (not junk). Coherence changed nothing here — identical to rules on all 332 photos.

Source: 332-photo holdout (test-holdout-ids-canonical.json), frozen 2026-08-25.

raw data (JSON) →