September 7, 2026 · 1 min read

Did the model change help?

Keep a short set of tasks from your own work: a few inputs, the answer you wanted, and what would be unacceptable. Ten is enough to start. They should be tasks you already do, not puzzles written to flatter a model.

Run the old and new model on the same set. Read the failures. A single average score hides the one case you cannot ship.

If you cannot name a task that improved, you do not have a reason to switch, even if a public benchmark moved.

All notes

Did the model change help? · sortNow Learn