Test Matrix — every stripe-accuracy test, one place

Two ad-hoc runs, the 8 researched candidates (Phase 0–4), and Alex's frontier rematch. Every done test links to a plain-English result page. DONE IN PROGRESS BLOCKED TODO · 2026-09-07

Notes save to this browser automatically. Export copies them to your clipboard (paste to Skippy) and downloads a JSON.
#TicketGroupTitleWhat it is · how it works · resultStatusBlocked reasonView resultOpen questionsData / resultYour notes
AH-1SAIL-148 Ad-hoc (already run)Multi-AI stripe-recovery test: Claude/Gemini/Nano-Banana/GPT-5.1 spot+draw missing stripes, rule-filtered Give four big AIs (Claude, Gemini, Nano Banana, GPT) the photos where our model missed a stripe. Ask each AI to draw the missing stripe. Score each drawn line against the human line, after we throw out any line that breaks our sail rules. RESULT: Claude recovered 78% of the missing stripes, Gemini 61%, Nano and GPT lower. The order did not change after the rule filter.DONE summary: https://other-ais-research.pages.dev/results/multi-ai-recovery.html
testing page: https://ai-stripe-recovery.pages.dev
Alex to rate the page (an LLM can't score itself). Fair rematch on untrained photos still owed (=AH-3).88 photos. Recovery of truly-missing stripes: Claude 78%, Gemini 61%, Nano Banana, GPT last. Ranking held after never-break rule filter; Nano lost 43 lines to geometry rules. Page: ai-stripe-recovery.pages.dev
AH-2SAIL-147 Ad-hoc (already run)Frontier-model drawing benchmark: traced lines vs v26/deep20 vs human labels (37 photos) Give one frontier AI (Fable/Opus class) full photos and ask it to draw all the stripes. Compare its lines to our two models and to the human labels on 37 photos. RESULT: it found 85% of stripes (our models find ~97%) and its lines were as accurate as deep20 when it committed, but it often declined to draw. Note: our models had trained on these photos; the AI had not, so this was not a fair fight.DONE summary: https://other-ais-research.pages.dev/results/frontier-bench.html Rematch on photos nobody trained on (=AH-3). Decline may be promptable.Found 85% of stripes (v26 97%, deep20 96%); per-line accuracy on par with deep20 when it committed; main failure = declining. Caveat: our models trained on these photos, frontier saw them cold.
AH-3SAIL-149 Alex's testFair frontier-AI rematch: rerun the big model WITH never-break rules + proper prompts, on photos it did NOT train on Repeat AH-2 the FAIR way: give the frontier AI our sail rules in the prompt, and use photos NONE of our models trained on. This is the rematch Alex asked for. EXPECT: shows whether a big AI, told the rules and tested cold, can match our trained model.BLOCKED Frontier model (Fable-5) quota exhausted; needs quota to run the fair rematch.summary: https://other-ais-research.pages.dev/results/frontier-rematch.html Which frontier model(s)? Which untrained photo set? Prompt wording for the rules. This is the test Alex explicitly asked for and I never ran.Fable-5 quota blocked the fair rematch. On the sample: v26 96.1% found / 0.225 false-per-photo, deep20 95.3% / 0.40; frontier unmeasured. Needs frontier quota.
AH-4SAIL-173 Alex's testCombined pipeline accuracy: our model + big-model rescue, step-by-step funnel Run the FULL pipeline end to end and show the funnel: how many real stripes our model finds alone, how many of the ones it MISSES the big model (Claude) recovers when asked, and the COMBINED final accuracy vs our model alone. Answers whether bolting on a big-model rescue actually makes a sailor's scan better.DONE summary: https://other-ais-research.pages.dev/results/combined-accuracy.html Running now (SAIL-173). Does adding a big-model rescue step to our pipeline raise final scan accuracy, and by how much?Combined pipeline: our model 68% -> +Claude rescue -> 93% recall on the hard photos (+25 pts), cost 0.24 junk/photo. Whole-holdout projection ~+3.3 pts (recall only).
0.1SAIL-150 Phase 0 (measure only)Per-stage ablation scoreboard on the 332 (finder alone / +rules / +SAM3 redraw) Run the 332 test photos through the pipeline three ways and score each: model alone, model + rules, model + rules + SAM3 redraw. Shows what each layer adds. RESULT: model alone finds nearly every stripe but draws extra junk (recall .998 / precision .902); adding rules cuts the junk hard (precision .978) for a small recall cost.DONE summary: https://other-ais-research.pages.dev/results/ablation.html Add the +SAM3-redraw column (not yet in the file).Frozen-332: raw recall .998/prec .902 -> +rules recall .958/prec .978 -> +coherence. Rules trade a little recall for big precision. File: gate0-ablation-scoreboard.json
0.2SAIL-151 Phase 0 (measure only)Label audit + blind random control: are our training labels under-traced? Take the 200 training labels our model most disagrees with, plus 100 random labels as a control, and have Alex mark each label good or bad. Tests whether our own training labels are wrong (skipping faint stripes). HOW: blind side-by-side gallery, no scores shown. EXPECT: if the top-100 are much worse than the random-100, bad labels are a real, fixable cause.IN PROGRESS summary: https://other-ais-research.pages.dev/results/label-audit.html
testing page: https://gate0-label-audit.pages.dev
Alex's bad-rate on top-100 vs random-100 base rate. This gates all of Phase 2.200-card blind gallery live (gate0-label-audit.pages.dev): 100 worst-disagreement + 100 random control. Early finding (SAIL-172): ~30 of 200 cards (~15%) have a long near-straight North Sails label chord that ignores the real stripe — bad labels the model correctly disagrees with. Awaiting Alex's full grades.
0.3SAIL-152 Phase 0 (measure only)Confidence-to-error stratification on the 332 For every stripe the model draws on the 332, join its confidence score to how wrong the line actually is. Tests whether low confidence predicts a bad line (so we can gate on it). RESULT: the least-confident fifth of detections match a real stripe only 27% of the time — so confidence is a usable quality signal.DONE summary: https://other-ais-research.pages.dev/results/confidence.html Per-band SAM3-redraw delta column still to add.1286 detections quintile-bucketed by YOLO confidence. Bottom quintile matches a real stripe only 27% of the time -> confidence DOES predict quality. File: gate0-confidence-stratification.json
0.4SAIL-153 Phase 0 (measure only)Curve-fit residual probe: robust low-order fit per stripe, correlate residual with error Fit a smooth low-order curve to each drawn stripe and measure how far the real points sit from that curve. Tests whether 'wiggliness' flags a bad line. EXPECT: a free per-line quality score, and it catches lines that grab a batten instead of the stripe.DONE summary: https://other-ais-research.pages.dev/results/curvefit.html Not run. Would give a free per-line quality score + second gate signal.Wiggle-vs-error correlation 0.62 (Spearman); a free red-flag quality signal layered on confidence. 1.2% batten-grabs.
0.5SAIL-154 Phase 0 (measure only)Candidate-pool recall check: are missed faint/bottom stripes present in SAM3-solo's 1816 candidates? Look at the 1,816 lines SAM3-solo proposes and check whether the faint/bottom stripes we currently MISS are hiding in that pool. Tests whether a filter could add missed stripes back, or only remove junk. EXPECT: decides if the junk-filter test (1.3) is worth building.DONE summary: https://other-ais-research.pages.dev/results/candidate-pool.html Not run. Decides whether a LUNA16 filter can ADD recall or only cut junk (gates 1.3).50% of missed stripes (29/58) ARE in SAM3's candidate pool -> a filter can recover up to half; other half needs a better proposer. GREENLIGHT 1.3 with a ceiling.
0.6SAIL-155 Phase 0 (measure only)Sail-outline registration probe: global sail fit to attack all three weaknesses Fit one outline of the whole sail, then check whether missed stripes fall inside the sail and false lines fall outside it. Tests whether one global sail shape can catch all three weak spots at once. EXPECT: a single geometry check that helps faint, bottom, and off-sail errors together.DONE summary: https://other-ais-research.pages.dev/results/sail-outline.html Not run.One sail-outline fit is strong for the MISS side (100% of bottom-band misses fall inside the sail) but weak as a junk filter (~36% of junk caught).
0.7SAIL-156 Phase 0 (measure only)Model soup of deep-slot-11/15/20 (uniform average, one holdout eval, recalibrate BatchNorm) Average the weights of three sibling models (deep-slot 11/15/20) into one 'soup' model and score it once on the 332. Tests whether free accuracy is available with no change to serving. EXPECT: a small accuracy bump at zero extra cost, or a clean negative.DONE summary: https://other-ais-research.pages.dev/results/model-soup.html Not run. Free accuracy at zero serving change if positive.Model soup = NO. Averaging v11/v15/v20 tanks precision (~50% vs 83-88%) via over-detection; not free accuracy.
0.8SAIL-157 Phase 0 (measure only)Single-valuedness audit of the 332 labels along luff-leech axis Check every label on the 332 for a stripe that doubles back on itself along the front-to-back axis. Tests whether the lane-detection model (CLRerNet, 3.2) can even be used. EXPECT: go/no-go — if labels double back, we need a different model head.DONE summary: https://other-ais-research.pages.dev/results/singlevalued.html Not run. CLRerNet go/no-go (gates 3.2).GREENLIGHT CLRerNet: only 1.28% of stripes double back, well under the go/no-go bar.
0.9SAIL-158 Phase 0 (measure only)Escalation stats for Nate's 'fewer-than-3-stripes' rule Count how often a real sail comes out of the pipeline with fewer than 3 stripes, and see what solo SAM3 finds on exactly those photos. Sizes Nate's 'call the big model only when we find <3 stripes' idea. RESULT: 0% come out short from the raw model, but 7.3% (24 of 329) end short after rules — so the trigger fires on those 24.DONE summary: https://other-ais-research.pages.dev/results/escalation.html What solo-SAM3 finds on exactly those 24 photos (the win size) still to measure.Frozen-332: 0% of sails come out of RAW model with <3 stripes, but 7.3% (24/329) end with <3 AFTER rules+coherence. So the trigger fires post-rules. File: gate0-escalation-stats.json
1.1SAIL-159 Phase 1 (1-day builds)TTA day over deep20 (multi-scale, bottom-crop, low threshold, custom polyline fusion) Run our existing model several ways on each photo (different zoom levels, a zoomed bottom crop, a lower confidence bar) and merge the results. No retraining. Tests whether we can squeeze out missed stripes for free. EXPECT: recovers some misses; if it recovers nothing, that is evidence the model never learned faint stripes (a label problem).DONE summary: https://other-ais-research.pages.dev/results/tta.html Not run. Null result = evidence for label root cause. Gated on Phase 0.TTA net-negative as-is: recovered 35 missed stripes but added 132 false lines (66 dup + 91 far). The weak recovery is itself evidence the miss is a LABEL problem, not the model.
1.2SAIL-140 Phase 1 (1-day builds)DeepLSD zero-shot boom finder + straightness veto Run a pretrained straight-line detector to find the boom (a long straight line) and to veto battens/rigging (straight) versus stripes (curved). No training. EXPECT: a free boom finder and a principled way to reject straight false lines. Needs some boom ground-truth first (Decision D1).DONE summary: https://other-ais-research.pages.dev/results/deeplsd-boom.html Needs boom ground truth on the 332 (Decision D1). Overlaps SAIL-140.1.2 = NO. Zero-shot line detection notices the boom 89% of the time but can't tell it from other straight things (rails, seams, dock lines) — picks the right one only ~35%. And straightness is a bad junk filter: catches 8% of fake lines while wrongly killing 13% of real stripes. Not worth building on.
1.3SAIL-134 Phase 1 (1-day builds)Junk-vs-stripe crop classifier (LUNA16 stage 2) on SAM3-solo's 1039 junk / 777 real Train a small yes/no classifier on SAM3-solo's 1,039 junk lines and 777 real lines to tell junk from stripe. Tests the 'over-propose then filter' pattern. EXPECT: cuts junk sharply; only worth it if test 0.5 shows real misses live in the pool.DONE summary: https://other-ais-research.pages.dev/results/junk-classifier.html Only if 0.5 shows misses live in the pool. Overlaps SAIL-134.Junk classifier AUC 0.89 from geometry alone; cuts total misses ~33% but costs ~1 extra junk line/photo at safe recall. Prototype further, don't ship as-is.
1.4SAIL-160 Phase 1 (1-day builds)Multi-prompt SAM3 redraw (points + loose box + mask prior) — retest a closed negative with the one untested fix Re-run the SAM3 redraw, but this time give it points, a loose box, AND a rough mask, instead of points alone. Re-tests a past failure with the one input we never tried. EXPECT: shows whether richer prompts fix the redraw; first confirm our SAM3 even accepts a mask.BLOCKED SAM3 mask-prior support CONFIRMED (unblocked technically); the points-only vs points+box+mask comparison needs a live GPU pod — all pods currently EXITED. Harness ready.summary: https://other-ais-research.pages.dev/results/multiprompt-sam3.html First verify our SAM3 accepts a mask prior at all. Run after 0.3's per-band table.1.4 PARTIAL: confirmed our SAM3 DOES accept the richer prompt (points + box + mask prior) — verified in the source, so the idea is testable. But the actual points-only vs richer-prompt comparison needs a GPU pod (all are currently down); the run harness is built and ready. Baseline being beaten: points-only SAM3 redraw = 36.4px err / 39.6% recall vs Deep20's 23px / 91.1%.
1.5SAIL-161 Phase 1 (1-day builds)Escalation gate v1: wire the fewer-than-3 rule to trigger solo SAM3 + 1.3 filter on flagged photos Wire the escalation gate: when a photo comes back with <3 stripes, automatically send it to solo SAM3 plus the junk filter. Builds what 0.9 sized. EXPECT: recovers stripes on the hard photos without slowing down the easy ones.DONE summary: https://other-ais-research.pages.dev/results/escalation-gate.html 0.9 sizes it; 1.3 supplies the filter.Escalation gate rescues 14/24 broken scans genuinely (6 are false rescues with a junk line). Broken-scan rate 7.3%->~3% if scored honestly. Worth building; gate on truth not count.
2.1SAIL-162 Phase 2 (data work)Worst-first relabel pass driven by the 0.2 audit ranking Feed the 0.2 audit ranking into the human relabel queue, worst labels first, and fix them. This is the root-cause fix every retrain inherits. BLOCKED until Alex finishes grading 0.2. EXPECT: a cleaner training set, so the next model learns the faint stripes it currently skips.BLOCKED Gated on 0.2 — needs Alex's label-audit grades first (the ranking drives which labels to relabel). Gated on 0.2 (Alex's grades). Alex sets batch size. The root-cause fix everything inherits.
2.2SAIL-163 Phase 2 (data work)Synthetic faint-stripe degradation + batten/rigging hard negatives Manufacture training examples for free: fade well-labelled stripes to look faint, and cut out battens/rigging as 'do not draw' examples. Tests whether cheap synthetic data teaches faint-stripe and anti-batten skills. EXPECT: more faint-stripe hits and fewer batten grabs, at zero labelling cost.TODO Fade well-labeled stripes to manufacture faint examples at zero labeling cost.
2.3SAIL-164 Phase 2 (data work)Pseudo-label the ~10k unlabeled uploads (only after 2.1 fixes the teacher) Auto-label the ~10,000 unlabelled uploads using the model, then train on them. BLOCKED until 2.1 fixes the teacher, or it copies our blind spot at scale. EXPECT: a big free data boost, but only safe after the labels are clean.BLOCKED Gated on 2.1 — pseudo-labelling the 10k uploads must wait until the teacher model is retrained on cleaned labels, or it copies our blind spot at scale. Must wait for 2.1 or it bakes the blind spot in harder.
3.1SAIL-165 Phase 3 (training bets)YOLO26-pose swap (NMS-free hypothesis; rides the next retrain) Retrain with the newer YOLO26-pose model instead of our current one. Tests whether its 'no-NMS' decoding stops adjacent stripes from suppressing each other. EXPECT: better separation of stacked parallel stripes; gains still capped by label quality.BLOCKED Blocked on a GPU pod — the training/build run couldn't reach a cloud pod from the local sandbox; needs to be run on RunPod.summary: https://other-ais-research.pages.dev/results/yolo26.html Bump pod ultralytics pin. Gains capped by label quality.
3.2SAIL-166 Phase 3 (training bets)CLRerNet fine-tune (rotate-90 converter) + ImageNet-init arm for license-clean promotion Fine-tune CLRerNet, a self-driving lane detector, by rotating photos 90° so stripes look like road lanes. Tests the closest-matching model in all of computer vision as a challenger finder. EXPECT: better faint/bottom-stripe recall than our model, if 0.8 passes.BLOCKED Blocked on a GPU pod — the training/build run couldn't reach a cloud pod from the local sandbox; needs to be run on RunPod.summary: https://other-ais-research.pages.dev/results/clrernet.html Gated on 0.8. Fallback = MapTR-style point-set head.
3.3SAIL-167 Phase 3 (training bets)DINOv3 linear probe at 1024px+ (grouping head only if probe passes) Freeze Meta's DINOv3 vision backbone and train a tiny probe to light up faint-stripe pixels, at high resolution. Tests whether it even 'sees' the paint our models miss. EXPECT: a go/no-go — if it sees faint stripes, build a full head; if not, drop it.BLOCKED DINOv3 weights are behind a click-through; needs a pod with the backbone downloaded.summary: https://other-ais-research.pages.dev/results/dinov3.html Low resolution makes a failure meaningless. Probe photos exclude the 332.DINOv3 backbone download needs click-through auth; only a CPU data audit ran. Needs a pod with the weights.
3.4SAIL-168 Phase 3 (training bets)PoseFix recipe for the SAM3 layer: retune redraw on deep20's real output distribution Retune the SAM3 redraw on our real model's output, including 'already correct' examples so it learns to leave good lines alone. Tests the known fix for refiners that always change something. EXPECT: redraw stops making good lines worse.BLOCKED Blocked on a GPU pod — the training/build run couldn't reach a cloud pod from the local sandbox; needs to be run on RunPod.summary: https://other-ais-research.pages.dev/results/posefix.html The literature's fix for refiners that always change something.
3.5SAIL-169 Phase 3 (training bets)clDice / skeleton-recall loss in the SAM3 fine-tune Add a training loss (clDice / skeleton-recall) that makes missed centre-line pixels expensive, in the SAM3 fine-tune. Only if SAM3 stays in production. EXPECT: fewer truncated faint stripe ends.BLOCKED Blocked on a GPU pod — the training/build run couldn't reach a cloud pod from the local sandbox; needs to be run on RunPod.summary: https://other-ais-research.pages.dev/results/cldice.html Only if SAM3 mask lane confirmed staying in production.
4.1SAIL-170 Phase 4 (system)Temporal burst agreement on Cape31 boom-cam bursts On the Cape31 boom-camera bursts, if a stripe shows in the frame before and after but not the middle one, re-prompt SAM3 there. Uses time, not averaging. EXPECT: recovers stripes that flicker out for one frame; only when the boat's trim is steady.TODO Gated on same-trim-state. Lives inside the Cape31 program.
4.2SAIL-171 Phase 4 (system)MM-Grounding-DINO fine-tune with explicit boom + batten classes Fine-tune MM-Grounding-DINO with explicit 'boom' and 'batten' classes to box them, feeding SAM3. BLOCKED on getting boom/batten box labels (Decision D1). EXPECT: named detections for boom and batten, if the labels exist.BLOCKED Needs a GPU pod — MM-GDINO (~900MB + backbone + text encoder) exceeds the safe local RAM budget at inference (caused a 6.5GB OOM before). Pod-run spec recorded.summary: https://other-ais-research.pages.dev/results/mmgdino.html Blocked on Decision D1's boom/batten boxes.4.2 BLOCKED — needs a GPU pod. MM-Grounding-DINO's ~900MB checkpoint plus backbone + text encoder exceeds the safe local memory budget at inference (this is what OOM'd the Mac before), so the zero-shot screen wasn't run locally. Checkpoint sizes verified without downloading. Pod-run spec is recorded for whoever runs it.

Below: the full research + test plan each matrix row links into.

Other AIs Beyond SAM3 — Researched Shortlist (2026-09-05)

From: Skippy (Nate's AI) · For: Alex (and his AI), Nate How this was produced: a 6-angle web research sweep covering lane-detection/vector-map models, thin-structure segmentation, the post-SAM3 open-vocab foundation stack, VLM pointing/grounding, pose-model upgrades, and adjacent domains plus technique tricks. The sweep surfaced 44 raw candidates; we deduped and ranked them to 8, then a separate agent adversarially fact-checked every finalist by opening the actual repo, paper, and license (several fetched checkpoints to confirm they download). A final completeness critic reviewed the whole sweep. Claims below carry the verifier's corrections rather than the researcher's optimism.

The question (Alex, 2026-09-04): "im convinced we can get better accuracy with using some kind of Ai and is there any other than SAM3 we should be looking at?"


The one-paragraph answer

No better "find faint painted stripes" brain sits on a shelf. The only real SAM3 successor as of Sept 2026 is SAM 3.1, a video-speed release with no accuracy gains, and 2026 chatbot/vision AIs still fail at tracing thin lines: three 2026 papers show they lose the line and jump to a nearby look-alike, which is exactly our batten problem. The closest architectural match to our problem in all of computer vision is lane detection from self-driving cars (faint painted line on a surface in, ordered point sequence out), and one proven, permissively licensed model, CLRerNet, is worth fine-tuning as a challenger finder. The two highest-value moves cost almost nothing and use no new model at all: a pretrained straight-line detector to find the boom and veto battens/rigging, and a label audit that finds our own under-traced training labels — the root cause behind the missed faint and bottom stripes.

The verified shortlist (rank = recommended order, after critic re-rank)

1. Label audit: find our under-traced training labels

2. CLRerNet, the lane-detection challenger finder

3. DeepLSD / ScaleLSD: boom finder plus batten/rigging veto, zero training

4. Test-time augmentation over deep20 (multi-scale, bottom-crop, low threshold)

5. DINOv3 dense backbone: faint-stripe evidence heatmap (linear probe first)

6. clDice / Skeleton Recall loss (only if SAM3 stays in the production path)

7. MM-Grounding-DINO: open-vocab boxes for boom and batten classes

8. YOLO26-pose, the drop-in finder swap

What got dismissed, and why it matters

Cross-cutting rules from the critique (adopt for every test above)

  1. Holdout hygiene: carve a dev slice from training data for tuning thresholds and margins; touch the frozen 332 only for final gates; apply the 73-content-dupe exclusion (exclude-holdout-dupes.mjs) to every holdout diff.
  2. Cheap data levers the sweep under-weighted: synthetic faint-stripe degradation (fade well-labeled stripes to manufacture faint training examples at zero labeling cost) and batten/rigging hard negatives. Pseudo-labeling the ~10k unlabeled uploads is real but must wait until the label audit fixes the teacher, or it bakes in exactly our blind spot.
  3. The boom-labeling budget is one decision: DeepLSD scoring, MM-GDINO classes, and a pose boom-class all want the same one-time boom annotations. Decide the target once; Alex's Cape31 correction pass is the natural vehicle (docs/cape31-program-plan.md step 4).
  4. Already answered internally: the critic asked whether SAM3 full-image concept prompting was ever tried as a recall source. It was: hop-eval run D (SAIL-132) is exactly that, and it found 1039 junk lines vs 777 real. The pending idea is the curvature gate (stripes sag; battens and shrouds stay straight; SAIL-134), which is also what makes DeepLSD's veto principled.

Recommended sequence (cheap to expensive, each gated on the last)

  1. Label audit plus random control (script only); SAIL ticket to follow.
  2. TTA day on deep20 (no training; doubles as a test of the label-root-cause hypothesis).
  3. DeepLSD zero-shot boom/veto screen (needs the boom-label decision).
  4. Single-valuedness audit of the 332 labels, then the CLRerNet fine-tune (the challenger-finder bet).
  5. DINOv3 high-res linear probe (one pod-day).
  6. YOLO26 swap when a retrain is scheduled anyway.
  7. clDice/skeleton-recall only if SAM3's production role is reconfirmed; MM-GDINO fine-tune only after boom/batten boxes exist.

Every eval number quotes its id file and freeze date, per house rules. Nothing auto-promotes.


Added 2026-09-05: How experts combine multiple models ("layering")

Second research sweep, four angles: production cascades, ensembling, foundation-model composition, and geometric/non-model layers. It surfaced 27 patterns, and every worth-adding claim was adversarially fact-checked against its primary sources; the corrections from that verification are baked into what follows.

The six rules the field follows

  1. A refiner only works on inputs like the ones it trained on. Cascade R-CNN's core finding: running the same model twice on its own output adds nothing, and running a refiner on inputs better than its training band makes them worse. This single rule explains both of our past negative results, the keypoint model redrawing its own zoomed crops and the early SAM3 refine variant. Both are the documented failure mode of the pattern done wrong.
  2. Split recall and precision into different models. The medical-imaging CAD pattern: stage 1 deliberately over-proposes and a small stage-2 classifier kills the junk; the LUNA16 lung challenge institutionalized this with candidate sensitivity around 98% before filtering. Under this rule our SAM3-solo result (1039 junk vs 777 real) reads as a healthy over-sensitive stage 1 that is missing its stage 2.
  3. Judge before you redo. A verifier that only scores, accepts, or rejects can never make a good output worse, while an ungated redrawer can. Gates need their own calibration check, because miscalibrated confidence routes wrongly in both directions. Google's shipped example: MediaPipe hand tracking re-runs its detector only when landmark confidence drops.
  4. Downstream layers cannot add recall. The finder's proposal set is a hard ceiling. Recall for faint/bottom stripes and booms must come from an upstream over-proposer or from a geometric prior that says where to look, never from more refinement.
  5. Ensemble gains come only from diversity. Re-running the same model adds zero, which is why our YOLO-on-its-own-crops failed while YOLO+SAM3 agreement works. And models trained on the same labels share blind spots, so agreement can be confidently wrong about the same faint stripe.
  6. Keep the committee offline; ship one model. Netflix's $1M prize stack was never deployed, and Google distills ensembles into single models. Every proposed layer must show its marginal number on the frozen 332 with the stage bypassed, or it doesn't earn a place.

Validation: our pipeline is already the textbook stack


Added 2026-09-05: THE TEST PLAN — every candidate, sequenced

Scope: all 8 shortlist candidates above, 5 layering additions, the gaps the completeness critic found, Nate's escalation idea, and Alex's frontier-model question. Nothing here runs until Nate/Alex say go. Proposed discipline: max 2 experiments in flight at once; each posts its result to the chat when it finishes; every final number is scored on the frozen 332 (test-holdout-ids-canonical.json, frozen 2026-08-25) with the 73 content-dupes excluded; all threshold tuning happens on a dev slice carved from training data and never on the 332; no model promotes without a human call.

Phase 0: measurement only, no training, no builds (about this week, under $5)

# Test What it answers Effort
0.1 Per-stage scoreboard: end-metric on the 332 with each existing layer bypassed (finder alone / +rules / +SAM3 redraw) The ablation table every later item is judged against ½ day, CPU, existing envelopes
0.2 Label audit + blind random control (shortlist #1); the audit script is already running as of this morning Whether under-tracing is real and how widespread: top-100 worst-scored labels vs a random-100 base rate, human-eyeballed script done; one human hour
0.3 Confidence-to-error stratification: join per-stripe YOLO confidence with line error on the 332, plus SAM3-redraw delta per confidence band Whether confidence predicts quality (enables gates and a customer confidence badge), and where redraw helps or hurts ½ day, inference-only join
0.4 Curve-fit residual probe (layering): robust low-order fit per stripe (Huber or leave-one-out at our ~10-point count, where classic RANSAC is statistically thin), correlate residual with error, count batten-grab flags A free per-line quality score and a second gate signal. Today stripeMetrics' Catmull-Rom passes through every raw point, so one bad point bends the numbers we sell 1 afternoon, CPU
0.5 Candidate-pool recall check: are the faint/bottom stripes we currently miss present among SAM3-solo's 1,816 candidates? Whether the LUNA16 filter path can add recall or only cut junk; decides 1.3 ½ day, data on disk
0.6 Sail-outline registration probe (layering, sports-field pattern): outlines on the 332 via the sail-edge spike method; count missed stripes inside their predicted height band and known FPs outside the sail Whether one global sail fit can attack all three weaknesses. Template stays loose; sails are soft, unlike soccer pitches 1 day
0.7 Model soup of deep-slot-11/15/20 (same init, same data, same sampling: the verified-correct trio, whereas v26+deep20 have different objectives): shape-assert, uniform-average, one holdout eval; recalibrate BatchNorm before calling a negative Free accuracy at zero serving change ½ day, 1 eval
0.8 Single-valuedness audit of the 332 labels along the luff-leech axis CLRerNet's go/no-go (3.2); failure reroutes to a MapTR-style point-set head 2 hours
0.9 Escalation stats for Nate's rule ("only run the big model when fewer than 3 stripes are found"): how often production returns fewer than 3 stripes, and what solo SAM3 finds on exactly those photos Sizes the win of the escalation gate before building it. This is the industry's confidence-gated cascade pattern (NoScope, MediaPipe) ½ day

Phase 1: one-day builds, no retraining (gated on Phase 0)

# Test Gate to proceed Effort
1.1 TTA day over deep20 (2–3 scales, zoomed bottom-third crop, lowered threshold; manual passes since Ultralytics' TTA flag doesn't run for pose; no horizontal flip; custom polyline fusion). A null result counts as evidence for the label root cause proceed regardless; judge on recovered misses vs new batten FPs 1 day
1.2 DeepLSD zero-shot boom finder and straightness veto (collinear-merge long segments; curvature margin mandatory) needs boom truth on the 332, so Decision D1 1 day
1.3 Junk-vs-stripe crop classifier (LUNA16 stage 2): train on SAM3-solo's 1039 junk and 777 real (photo-level split, 332 untouched), measure ROC only if 0.5 shows the misses live in the pool; otherwise its value is FP reduction alone 1 day, minutes of GPU
1.4 Multi-prompt SAM3 redraw (the SAMRefiner idea: points, a loose box, and a mask prior mined from the polyline). Framed plainly as re-testing a closed negative with the one untested fix; first verify our SAM3 accepts a mask prior at all run after 0.3's per-band table exists 1 day plus GPU hours
1.5 Escalation gate v1: wire Nate's fewer-than-3-stripes rule (plus low-confidence routing if 0.3 shows signal) to trigger solo SAM3 plus the 1.3 filter only on flagged photos 0.9 sizes it; 1.3 supplies the filter 1–2 days

Phase 2: data work (gated on 0.2's audit results)

Phase 3: training bets (each is one run plus dual-ruler proof; gated on Phase 2)

Phase 4: system-level

Not doing, and why (one line each)

VLM line-tracing (three 2026 papers show them losing the line to a nearby look-alike, exactly battens); hosted APIs like DINO-X and Gemini trajectories (photos leave our infra and nothing can fine-tune); WBF-style averaging fusion (selection beats fusion while confidences are uncalibrated); a learned combiner (the Netflix lesson; our rules layer already plays that role, debuggably); deep ensembles in serving (artifact count, though the snapshot variant is fine offline for committee QC); top-down pose models (stripe crops overlap near-fatally); migrating to SAM 3.1 (a video-speed release).

Answers folded in

Decisions needed before the gated items