← Back to writing

LLM service tiers: which AI features belong on flex vs priority

Short answer: the big model APIs now sell the same model at different speeds and prices. Flex (about half price, slow, can be refused), standard, and priority (a premium for faster, steadier latency). Most SaaS teams send every call at standard. Map each AI feature to a tier based on who is waiting for the answer, and you can cut a large share of the bill without touching quality.

This post is the decision table we use when we review a client's LLM spend: what each tier actually promises, which features belong on which tier, and the fallback code you need so a cheap tier never becomes an outage.

What the service tiers actually are

A service tier is a request-level setting, not a different model. You send the same prompt to the same model and pass a service_tier value. The provider decides how fast your request gets compute, and bills you accordingly. As of October 2026 the main options look like this:

Flex: half price, best effort

OpenAI's flex processing (service_tier: "flex") is billed at Batch API rates, but you call it synchronously like a normal request. The trade is slower responses and occasional refusals. When there is no capacity, OpenAI returns a 429 "Resource Unavailable" error, and per its flex processing guide you are not charged for that request. The guide's examples raise the client timeout to 15 minutes, because the SDK default is 10.

Google's Gemini API has the same idea. Its Flex inference tier is 50% off standard, targets responses in 1 to 15 minutes, and runs on capacity that can be preempted when standard traffic spikes. Expect a 503 or 429 when it is full. Google is explicit that a Flex request is not automatically upgraded to Standard, so the retry or fallback is your job.

Standard: the default

This is what you get when you pass nothing. It has normal rate limits, normal latency and list price. It is the right tier for a lot of traffic, but it is rarely the right tier for all of it.

Priority: pay more for steadier latency

Priority tiers put your request at the front of the queue. Gemini's Priority inference costs 75 to 100% more than standard. When you go over your priority limit, the overflow is downgraded to Standard and billed at the standard rate instead of failing. A response header tells you which tier actually served the call. OpenAI sells the equivalent as service_tier: "priority" (marketed as Fast mode since mid-2026) at a premium over standard. On Anthropic's API, Priority Tier is tied to capacity commitments that are no longer sold to new customers. New teams there choose between standard and the Batch API, and service_tier accepts auto or standard_only (see Anthropic's service tiers page).

Prices and model lists change often. Treat the numbers above as a snapshot and check each provider's pricing page before you lock in a budget.

Map each feature to a tier by who is waiting

The useful question is not "how important is this feature", it is "who is waiting for this response, and for how long". That one question sorts most of a SaaS product's LLM calls.

A user is watching a spinner: standard, sometimes priority

Chat replies, inline autocomplete, an "explain this chart" button. A user is looking at the screen, so a 4-minute answer is a broken feature. Keep these on standard. Move them to priority only when you have measured that standard-tier tail latency is hurting you. For example, your p95 time to first token keeps blowing through a target you have committed to customers. Priority is a fix for a measured latency problem, not a default. If latency is the real complaint, read our breakdown of what drives time to first token first. Prompt size and output length usually matter more than the tier.

A user will come back later: flex

Weekly account summaries, "generate a report and email me", lead enrichment after a form submit, tagging a support ticket that nobody will open for an hour. The user has already left the screen. These calls can wait minutes, and they are often the bulk of token volume in a B2B product. This is the easiest money on the table. You make the same call, add one parameter and a longer timeout, and pay half.

Nobody is waiting: batch

Backfills, re-embedding a corpus, nightly evals, re-classifying a year of tickets. If the job can take up to 24 hours, the Batch API gives a similar discount and its own separate quota, so it does not eat into your live rate limits. We covered when batch beats synchronous calls in our batch API cost guide. Flex sits between the two. It costs the same, but you write ordinary request code, which is much simpler for jobs that run inside an existing worker.

A worked example

Take an illustrative B2B SaaS product spending $10,000 a month on one model at standard rates, with this split: 35% interactive chat, 45% background jobs triggered by user actions (summaries, enrichment, classification), 20% scheduled work (nightly digests, evals). Moving the background 45% to flex at half price saves about $2,250 a month. Moving the scheduled 20% to batch saves about $1,000 more. That is roughly a third off the bill with no prompt changes and no model downgrade. Your split will differ, so measure it per feature first. Our post on LLM cost attribution per feature shows how to tag calls so this table is built from real data instead of guesses.

The fallback code that makes flex safe

Flex fails in a new way: the provider can turn a request down. If your worker treats that as a hard error, a cheap tier becomes a reliability incident. The pattern we ship has four parts.

  1. Set a long client timeout on flex calls only. Fifteen minutes is a sane starting point. Keep interactive calls on a short timeout so they still fail fast.
  2. On a capacity error (429 or 503 from the flex tier), retry on flex with exponential backoff and jitter. On OpenAI the refused attempts are not billed, so waiting costs you time, not money.
  3. Give every job a deadline. If the job has not succeeded on flex by its deadline (say, 30 minutes for an emailed report), retry once on standard and accept the full price for that one call.
  4. Log the tier that actually served each response. OpenAI and Anthropic echo a service_tier field in the response, and Gemini sets a response header. Without this you cannot tell whether your savings are real or whether half your flex traffic is quietly escalating to standard.

Retries need care so they do not stack up. The same backoff rules from our guide to handling 429 rate limits apply here. Also make the job idempotent, so a retry that lands after a slow success does not run the work twice.

What to watch after the switch

Track three numbers per feature for two weeks: the share of flex calls that fell back to standard, p95 job completion time, and the effective cost per job. If the fallback rate climbs above about 20%, flex capacity for that model is tight at your traffic hours. Either move the job to batch or schedule it off-peak. If completion time breaks a promise you made to users ("your report arrives within 15 minutes"), that feature belongs on standard.

When not to bother

Tiering is plumbing work, and it does not always pay back. If your LLM bill is a few hundred dollars a month, the engineering hour costs more than the saving. If almost all your traffic is interactive chat, there is little to move. And if your provider does not offer flex on the model you depend on, a cheaper model on standard may beat the same model on a discount tier. Run the numbers in our AI cost calculator before you start.

FAQ

Does flex give worse answers than standard?

No. It is the same model with the same parameters. Only scheduling and price change. If your quality evals move after switching, look for a different cause, such as a timeout cutting off long outputs.

Can I use flex and prompt caching together?

On OpenAI, yes. The flex guide says prompt caching discounts apply on top of flex pricing, so a long shared system prompt gets cheaper twice.

Is priority worth it for a chatbot?

Only if you have measured a tail-latency problem that standard cannot meet. Shorter prompts, streaming and a smaller model for simple turns usually fix perceived speed for far less money.

How long does it take to roll this out?

For a codebase with one LLM client wrapper, adding a tier option, the fallback logic and per-tier logging is typically two to four days of work, plus two weeks of observation. If you would rather hand it to a team that has done it before, our LLM engineers can scope it against your actual traffic.

Get shipped

Rather we just build it?

Book a free scoping call and we'll ship your production-safe AI feature this week.