{
 "_meta": {
  "date": "2026-09-06",
  "photos": 37,
  "photoSource": "training-data-mixed (holdout-excluded; NOTE: v26/deep20 saw these photos in training, frontier saw them cold \u2014 asymmetry favors the incumbents)",
  "tracer": "claude-fable-5 vision agents, draw-render-check loop, up to 2 refinements",
  "scoring": "hop-eval v2 semantics: 60-pt densify, 3%-diag match cap, coverage=human->line, overhang=line->human"
 },
 "perSource": {
  "frontier": {
   "stripesFoundRate": 0.854,
   "falsePerPhoto": 0.486,
   "coverageMedPct": 0.226,
   "overhangMedPct": 0.247,
   "cleanPhotos": "21/37"
  },
  "v26": {
   "stripesFoundRate": 0.967,
   "falsePerPhoto": 0.297,
   "coverageMedPct": 0.159,
   "overhangMedPct": 0.157,
   "cleanPhotos": "29/37"
  },
  "deep20": {
   "stripesFoundRate": 0.959,
   "falsePerPhoto": 0.405,
   "coverageMedPct": 0.21,
   "overhangMedPct": 0.21,
   "cleanPhotos": "31/37"
  }
 },
 "verdict": "Not the scanner today; per-line accuracy on par with deep20 when it commits; failure mode = declining to trace (12/37 photos under-counted vs label). Greenlights judge/label-factory/rescue roles; fair rematch belongs on photos no model trained on, with a commit-forcing prompt."
}