August 26, 2026 · 1 min read
Three questions before you pay for another model
How long does a typical request take, including retrieval and any tool calls, not just the model? If a person is waiting, a second and eight seconds are different products.
What does a day of your real traffic cost? Price a prompt you actually send, at the length you actually send, with the retries you actually have. A price per million tokens is not a budget.
How often does a person have to correct the output? A cheaper model that needs a full rewrite is not cheaper. A faster model that is wrong in a way you cannot spot is not faster in the way that matters.