Weight-averaging our three sibling checkpoints sounded like a zero-cost accuracy bump — same serving setup, no retraining. It isn't: the averaged model gets confused and draws far too many candidate lines, most of them wrong. This path is closed; don't spend more time on model soups from these checkpoints.
| Model | Recall | Precision | Avg line error |
|---|---|---|---|
| v11 | 99.8% | 85.5% | 0.0072 |
| v15 (best single sibling) | 99.9% | 83.2% | 0.0063 |
| v20 | 99.8% | 87.7% | 0.0069 |
| Soup (v11+v15+v20 averaged) | 100.0% | 49.7% | 0.0105 |
A "model soup" averages the weights of several sibling models into one, hoping to get the best of all three for free. Here it matched or slightly beat recall, but precision collapsed — the soup draws roughly twice as many false lines as any single sibling.
Source: 332-photo holdout, frozen 2026-08-25. Raw model output, no production rules filter applied.