Alla artiklar
6 oktober 2026

Choosing an LLM for a product feature: quality, latency and cost tradeoffs

The right model for a feature is set by its quality bar, latency budget and cost ceiling, not by which model ranks highest overall.

Team in a bright office: a man posts sticky notes on a wall while colleagues work at laptops

Pick the model that clears a feature's quality bar inside its latency budget and cost ceiling. Not whichever model tops a general leaderboard. Every major vendor's own documentation frames the choice the same way: capability, speed and cost pull against each other, and picking a model means picking a point on that tradeoff.

What it is and why it matters

Capability, speed and cost are not three separate decisions. A model that clears your quality bar at twice the latency, or at ten times the cost, is usually the wrong one, even though it is technically "better." Anthropic's own model selection guidance tells builders to weigh capability, speed and cost together before picking a starting model. OpenAI's model selection framework reduces the same decision to four questions: how often the workflow runs, how fast the result needs to come back, how the output gets used, and how much quality the task actually demands. OpenAI states the rule plainly: keep the lightest setting that meets your quality bar.

There are two reasonable ways to start. The efficiency first path begins with a small, fast, cheap model, tests it against real prompts, and upgrades only where it fails on a specific capability. That fits prototyping, tight latency requirements, cost sensitive products, and high volume tasks like classification or extraction. The capability first path does the opposite. Start with the strongest available model, optimize the prompt for it, then step down to something cheaper over time once the feature is proven. This fits complex reasoning, technical or scientific work, and any case where getting the answer wrong costs more than the extra tokens would.

A feature built on retrieval-augmented generation shows why the model is only half the decision. When the retrieval step surfaces the right passage, a smaller model can often answer well. When retrieval is weak, even a frontier model will produce a confident, wrong answer. Model choice and retrieval quality are separate levers, and testing them separately is the only way to know which one is actually limiting the feature.

How it works in practice

Most teams change three things, usually in this order: the effort or reasoning depth setting, the model itself, and the billing mode the request runs under.

Effort is the smallest lever. Several vendors expose a parameter, often called effort or reasoning depth, that trades intelligence for latency and cost inside a single model, without switching models at all. Anthropic recommends tuning effort before switching models, since it's a smaller and more reversible change. OpenAI pairs its models with reasoning effort tiers the same way: a low setting for simple extraction and scoped problems, a medium setting for complex technical work with expected revisions, and an extra high setting reserved for demanding analysis with exacting requirements.

Billing mode is a second, independent axis on top of whichever model gets picked. OpenAI's pricing documentation, fetched October 6, 2026, lists batch processing at roughly half the standard per token rate, in exchange for latency measured in hours rather than seconds. A fast tier runs roughly double the standard rate for the lowest latency. Google's Gemini API optimization guide documents the same pattern across four tiers: standard runs at full price with latency of seconds to minutes, flex gives a 50 percent discount for a 1 to 15 minute target latency and can get preempted during traffic spikes, priority costs 75 to 100 percent more than standard for guaranteed seconds level latency, and batch gives a 50 percent discount for latency up to 24 hours with automatic retries.

The price spread between a flagship model and a smaller model in the same family can be large. OpenAI's pricing page shows an example flagship model priced at 10 USD per million input tokens and 50 USD per million output tokens on the standard tier, against a smaller model in the same lineup at 0.10 USD and 0.50 USD. A 100 times difference on both input and output rates. That spread is the economic argument for routing requests to the cheapest model that can still do the job, rather than sending every request to the same model by default.

Context caching is a third cost lever that has nothing to do with which model you pick. Google's optimization docs document a 90 percent discount on cached tokens when a large, repeated initial context, such as a long system prompt or a reference document, gets reused across requests, along with a faster time to first token on the cache hit.

Tradeoffs and edge cases

Multi-model routing is how teams apply all of this at once: send most requests to a cheap, fast model by default, and escalate only the ones that need it to a more capable model. Anthropic describes two versions of this, an executor that escalates hard decisions to an advisor model, and an orchestrator that delegates bulk work to lower cost worker models. A feature that calls out to MCP to reach tools or a bigger model mid task is one common way this routing gets implemented in an agentic product.

Before committing to a model, build an eval set from your own prompts and your own data. Every vendor's guidance converges on the same practice: test candidate models against real inputs, compare accuracy, response quality and how each one handles edge cases, then weigh the cost difference against what that accuracy gain is actually worth to the feature. A leaderboard score, or a vendor's own marketing claim, won't tell you that. Quality on a general benchmark doesn't always predict quality on your specific task.

Model choice also isn't the same problem as safety. For agentic or high autonomy features, guardrails and approval gates around what the feature is allowed to do usually matter more to the product's safety than which model sits behind it. A more capable model doesn't substitute for a permission boundary, an approval step before an irreversible action, or an audit log of what the feature actually did.

The most common mistake is defaulting to the most capable, most expensive model for every call in a feature, including the narrow, well scoped parts of it. Classification, extraction and short lookups rarely need a frontier model's reasoning depth. Routing that traffic to a cheaper model usually costs nothing in quality while cutting the bill by a wide margin.

Frequently asked questions

Should I always start with the cheapest model?

For most product features, yes. Start with the smallest or cheapest model that could plausibly do the job, test it against real prompts, and upgrade only where it misses a specific capability. Reserve a capability first start for tasks where accuracy outweighs cost from the outset, such as complex reasoning or long horizon agentic work.

What is the difference between changing models and changing effort?

Effort, sometimes called reasoning depth, is a setting inside a single model that trades intelligence for latency and cost without switching models at all. Vendors recommend tuning effort first, since it's a smaller and more reversible change than swapping the underlying model.

Does a cheaper model always mean worse quality?

Not for every task. A smaller model can match a flagship model on narrow, well scoped work like extraction or classification, and the only way to know is to run both against your own eval set rather than assume quality scales with price.

How much can batching or caching actually save?

Batch processing commonly runs at roughly half the standard per token rate in exchange for latency measured in hours instead of seconds. Reusing a large repeated context through caching can cut costs further still: one vendor documents a 90 percent discount on cached tokens, plus a lower time to first token.

A team scoping a new AI feature can apply the same test before writing any product copy about it: try the smallest model that could plausibly work, and only reach for the biggest one once a real gap shows up in the evals. Kallos Labs can act as that second opinion on where to draw the line; reach out through the AI automation page if that is useful.