September 7, 2026 · 1 min read
Did the model change help?
Keep a short set of tasks from your own work: a few inputs, the answer you wanted, and what would be unacceptable. Ten is enough to start. They should be tasks you already do, not puzzles written to flatter a model.
Run the old and new model on the same set. Read the failures. A single average score hides the one case you cannot ship.
If you cannot name a task that improved, you do not have a reason to switch, even if a public benchmark moved.