Idempotent LLM jobs: stop paying twice for the same AI call
Short answer: background AI jobs get retried by queues, workers, webhooks and impatient users, and every retry is a paid LLM call. Give each job a stable idempotency key, record the result before you acknowledge the job, and check that record before calling the model. This stops duplicate spend and duplicate side effects like the same summary emailed twice.
Most teams find this problem on the invoice or in a support ticket, not in code review. This post covers where the duplicate calls come from, how to design the key, and the storage pattern we use so a retried job returns the saved answer instead of paying for a new one.
Where duplicate LLM calls come from
In a normal web request, a duplicate call is rare and cheap. Background AI jobs are different. They are slow (seconds to minutes), they run inside systems that are built to retry, and each attempt can cost real money when the prompt carries a long document. The usual sources:
Queue redelivery
SQS, RabbitMQ, Pub/Sub and most job runners deliver at least once. If a worker takes longer than the visibility timeout, or crashes after the model call but before it acknowledges the message, the job runs again. An LLM call that takes 90 seconds against a 60-second visibility timeout is a duplicate on every single run.
Client and SDK retries
Official SDKs retry on timeouts and some 5xx errors by default. If the provider finished generating but the connection dropped before you got the response, you pay for the first generation and then again for the retry. The more you raise timeouts for slow models, the more this matters. We cover the timeout side in our guide to LLM timeouts and retries.
Webhooks and double clicks
Stripe, HubSpot, GitHub and most other webhook senders redeliver when your endpoint is slow to return a 2xx. A user who clicks "Generate summary" and sees nothing for eight seconds clicks it again. Each one starts a fresh job unless something stops it.
Agent loops
An agent that re-plans after a tool error often calls the same tool with the same input again. If that tool is itself an LLM call (a summarizer, a classifier), the agent pays twice for an answer it already had. Loops like this are one of the drivers behind runaway agent token costs.
How much this adds up to depends on your stack, and you can measure it in an afternoon. Log a hash of (job type, entity ID, prompt version, input) for every model call for a week, then count hashes that appear more than once inside the same job. Any repeat you find is money spent on an answer the system already had, and often a duplicate side effect, such as the same email sent to a customer twice.
Designing the idempotency key
The key answers one question: "is this the same unit of work I already did?" Get it wrong in one direction and you still pay twice. Get it wrong in the other and you return a stale answer for new input.
Build it from the business action, not the request
A random UUID made at enqueue time protects against queue redelivery but not against the double click, because each click makes a new UUID. A better key describes the work itself, for example summarize:ticket:48213:v3, where v3 is the ticket's revision number. Two clicks on the same ticket revision produce the same key. A new customer reply bumps the revision and correctly produces a new key.
Include what changes the answer
If you change the prompt template or the model, old results may no longer be valid. Add a short prompt version to the key (summarize:ticket:48213:v3:p7) so a prompt deploy naturally invalidates old entries. Do not include things that vary without changing the answer, like a request timestamp. That defeats the purpose.
Idempotency is not caching
An idempotency record says "this exact job already ran, here is its result". A cache says "a similar question was answered before, reuse it". Both save money, but they have different correctness rules. Semantic caching can return a near match, which is fine for some FAQ answers and wrong for a specific customer's ticket. Keep them separate. If you want the caching side too, see our semantic caching guide.
The storage pattern
You need one table (or a Redis hash) keyed by the idempotency key, with a status, the result, a lock expiry and a cost field. The flow for each job:
- Try to insert a row with status
runningand a lock that expires a bit after your maximum expected runtime. Use an atomic insert-if-absent (INSERT ... ON CONFLICT DO NOTHINGin Postgres,SET NXin Redis). - If the insert fails and the existing row is
done, return the stored result. No model call. - If the existing row is
runningand its lock has not expired, another worker has the job. Requeue with a delay rather than calling the model in parallel. - If the lock has expired, the previous worker probably died. Take over the row and run the job.
- After the model responds, write the result and set status to
donein the same transaction. Only then acknowledge the queue message or return the webhook 2xx. - Run side effects (sending the email, posting to Slack) as separate steps that check their own key, such as
email:ticket:48213:v3. That way a crash between "model done" and "email sent" retries the email without paying for the model again.
Step 5 is the one most teams get backwards. Acknowledge first and you lose work on a crash. Save the result first and the worst case is a cheap duplicate acknowledgement.
How long to keep records
Keep them for at least the longest retry window in your system: the queue's maximum receive count times the visibility timeout, plus your webhook senders' retry window. Many webhook providers retry for a day or more. A 7-day retention covers most stacks. Store the token counts and cost with each row, because this table becomes a good source for per-feature cost attribution at no extra effort.
Measure the savings
Add a counter for "returned stored result instead of calling the model", tagged by feature. In the first week this tells you how many duplicate calls you were paying for. After that it is a regression alarm: a sudden jump usually means a worker started timing out and the queue is redelivering again.
Where this fits in your architecture
Idempotency matters most for work you have already moved off the request path. If you are still deciding which AI features should run synchronously, in a queue, or ahead of time, start with our sync, async and precompute patterns, then add the key at the queue boundary. The whole change is usually a day or two for one job type: one table, one wrapper around your model client, and a key builder per job.
FAQ
Do OpenAI or Anthropic deduplicate requests for me?
Do not count on it. Treat every completed generation request as billed, and build deduplication on your side where you control the key and the stored result.
Should interactive chat use idempotency keys too?
Usually not for the model call itself, since each chat turn is new input. It is still worth a short-lived key on the send button to absorb double submits.
What if two different jobs produce the same key by accident?
That is a key design bug, and it returns the wrong answer, which is worse than paying twice. Include the tenant ID and the entity ID in every key, and add a test that two tenants with the same entity ID never collide.
Can I just lower SDK retries to zero?
You will trade duplicate spend for failed jobs. Keep retries for real transient errors and let the idempotency record make them safe. If you want a senior team to wire this into your job system, our AI developers do this as a standard task.
Rather we just build it?
Book a free scoping call and we'll ship your production-safe AI feature this week.