Your LLM endpoint sees a mix of work. One-line lookups, routine completions, and every so often a problem that needs a big model, all landing on the same API. Point all of it at your best model and you pay top rates for the easy requests, which are most