← Back to writing

Sync, async, or precompute: where an AI feature should run

Answer first: an AI feature in an existing SaaS product can run in one of three places. Synchronously, inside the request the user is waiting on. Asynchronously, as a job the user is told about when it finishes. Or precomputed, at write time, so the user-facing read is a plain database query. Most teams pick synchronous because it is the shortest path to a demo, then discover that a 6 to 20 second model call sitting inside a page load is the reason the feature feels broken. The right choice depends on three numbers: how long the model call takes at p95, how often the same input is asked about, and how fresh the answer has to be.

  • Synchronous fits interactive, one-off questions where the input is unknown until the user types it: chat over a document, rewrite this paragraph, explain this error.
  • Asynchronous fits anything that takes longer than about 5 seconds or touches many records: summarize this 400-page contract, classify every ticket in the queue, draft a weekly report.
  • Precompute fits anything that will be read many times but written once: deal summaries, ticket tags, account health notes, search embeddings.

Why the default synchronous call hurts

The demo pattern is simple. A user clicks a button, the server builds a prompt, calls the model, waits, and returns the answer in the same HTTP response. It works on stage because the demo input is short and the room is patient.

In production three things go wrong at once. First, latency. A single model call that reads 30k tokens of context and writes 800 tokens of output routinely takes 8 to 15 seconds, and the p95 is worse than the median because provider latency is bursty. Second, failure handling. If the model times out, the user's page times out with it, so a provider incident becomes your incident. Third, cost concentration. Every page view that triggers the call pays for it, even when the underlying record has not changed since the last view.

The timeout and retry rules for LLM calls in SaaS features cover the second problem. The first and third are architectural, and no retry policy fixes them.

The three patterns, and the numbers that pick between them

Pattern 1: synchronous, with streaming

Keep the call inside the request only when the input is genuinely unknown until the user acts and the user expects to wait. Even then, stream. A streamed response that shows its first token in under a second feels responsive at a total duration that would feel broken as a single blocking spinner. We covered the mechanics in how to stream LLM responses inside a SaaS feature.

Hard rules for this pattern: cap the context you send so the call cannot balloon with a large record, set a client-visible timeout that is shorter than your load balancer's, and have a fallback state that is not a blank error.

Pattern 2: asynchronous job with a status

If p95 latency is above roughly 5 seconds, or the work spans more than a handful of records, move it to a job. The request enqueues work and returns immediately with a job id. The UI shows an in-progress state, and the result arrives through polling, a websocket, or a notification. This is the same queue you already use for exports and email, so it costs nothing new in infrastructure.

The payoff is that model latency stops being user-facing latency. A 40-second batch classification is fine when the user is not staring at a spinner. Retries, rate-limit backoff, and idempotency keys all live in the worker, where they belong. Failure becomes a job state the user can retry, not a 504.

The cost is product work. Someone has to design the pending state, the done notification, and what happens if the user edits the source record while the job runs. Teams skip this pattern because it looks like more work, then spend longer fighting timeouts than the job would have taken to build.

Pattern 3: precompute at write time

If the same output will be read many times, generate it when the input changes, not when it is read. A CRM deal summary is a good example. Activity lands on the deal a few times a day. The summary is viewed by several people, several times each, and shown in list views and digests. Computing it on every view means paying for the same tokens dozens of times and making every reader wait. Computing it once per activity, in a worker triggered by the write, makes the read a database lookup that returns in milliseconds.

The arithmetic is simple. Suppose a record changes twice a day and is viewed twenty times a day. Synchronous generation costs twenty model calls. Precompute costs two. At 30k input tokens per call, that is the difference between 600k and 60k tokens per record per day. Multiply by the number of active records and it is usually the single largest line in the monthly LLM cost estimate for the feature.

Freshness is the tradeoff. A precomputed answer is only as current as the last write that triggered it. For most summaries, tags, and scores this is fine, because the underlying data is what changed and the model output just lags it by seconds. For anything where the user's question is the input, precompute does not apply.

A decision table you can apply in a planning meeting

Ask three questions about the feature on the whiteboard.

  • Is the input known before the user asks? If no, the feature is synchronous or asynchronous. If yes, precompute is on the table.
  • What is the p95 model latency with real-sized context? Under about 3 seconds, synchronous with streaming is acceptable. Between 3 and 8 seconds, streaming makes it survivable but a job is safer. Above 8 seconds, use a job.
  • What is the read-to-write ratio? If a generated output is read more than two or three times per change, precompute pays for itself in cost and latency immediately.

Many real features are hybrids. A support copilot might precompute the ticket summary and suggested tags on every ticket update, then answer the agent's free-form question synchronously with streaming, using the precomputed summary as context so the live call is small and fast. That split is what makes the live part feel instant: the expensive reading happened earlier, in the background.

What this means for the first ship

Pick the pattern before writing the prompt. It changes the data model, because precompute needs a column or table for the stored output and a version stamp for which prompt produced it. It changes the API, because async needs a job resource. And it changes the feasibility spike you run before committing to a date, because the latency number you measure in the spike is the one that decides the pattern.

When we scope an AI feature as a single plain-English task, the pattern decision is usually the first thing the engineer writes back, before any code. It is cheap to decide up front and expensive to change after launch, because moving a synchronous feature to precompute means backfilling every existing record and rewriting the read path.

If you are unsure, the safe default for anything that is not a live conversation is the job pattern. It is the easiest to change later, because a job can be triggered by a write (becoming precompute) or by a user action (staying async) with no change to the worker itself.

Should I always stream synchronous LLM responses?

Yes, when the response is longer than a sentence. Streaming does not reduce total latency, but it reduces time to first visible output to under a second, which is the number users judge. For short classifications or a single score there is nothing to stream, so just make the call fast by sending less context.

How do I keep precomputed outputs from going stale?

Trigger regeneration from the same event that changes the source record, and store the prompt version and model name next to the output. When you change the prompt, you can backfill in the background and know which records still carry the old version. The prompt versioning pattern for production AI features covers the bookkeeping.

Can I start synchronous and move to a job later?

You can, but budget for it. The move requires a job resource, a pending UI state, and a notification path, none of which exist in the synchronous version. If the model call might ever exceed a few seconds on real data, build the job version first. It is roughly a day of extra work up front and saves weeks of timeout debugging.

Get shipped

Rather we just build it?

Book a free scoping call and we'll ship your production-safe AI feature this week.