← Back to writing

Deleting customer data from an AI feature: the full map

Short answer: when a customer asks you to delete their data, deleting the row in your main database is not enough once you have shipped an AI feature. Their data now also lives in prompt logs, vector indexes, caches, eval datasets, fine-tuning files and sometimes in objects stored at your model provider. Map every one of those places before launch, key each copy to a tenant or user ID, and run deletion as one tracked job that fans out to all of them.

Most SaaS teams discover this the hard way. An enterprise customer churns, their contract says data is deleted within 30 days, and someone asks the engineering team to confirm. The app database is clean in an hour. Then someone remembers the support copilot indexed every ticket into a vector store, the observability tool kept full prompts, and three of the customer's tickets were copied into the regression eval set last quarter. Nobody can say with confidence that everything is gone.

This post is the map we use to avoid that situation, plus the deletion job design that makes the answer provable.

Why AI features multiply copies of customer data

A classic CRUD feature stores customer data in one or two places: the primary database and its backups. An AI feature adds a pipeline, and every stage of that pipeline tends to keep its own copy for good engineering reasons. Logs help you debug. Embeddings make retrieval fast. Caches cut cost. Eval sets catch regressions. Each is reasonable on its own. Together they mean a single support ticket can exist in eight places.

The legal side makes this matter. Under GDPR Article 17, a person can ask a controller to erase their personal data, and the controller has to do it "without undue delay" when one of the grounds applies. Article 12 gives you one month to respond, extendable in some cases. US state privacy laws such as the CCPA carry their own deletion rights, and many B2B contracts add a fixed deletion window on termination. Whatever the source of the obligation, the engineering problem is the same: you need to find every copy.

The AI data map: every place a customer's data can end up

Walk your feature end to end and write down each store below that applies. For each one, note the key you would use to find a customer's records and whether you can delete by that key today.

1. Prompt and response logs

Your own request logs, your LLM gateway, and any tracing or observability tool you send spans to. These usually hold the full prompt, which includes whatever retrieved context you stuffed in. If the logs are keyed only by request ID, deleting one customer means a full scan. Add tenant ID and user ID as indexed fields on every log line, and set a retention period short enough that most deletion requests resolve themselves by expiry. If you strip personal data before it reaches the model at all, as in our PII redaction gateway pattern, the log problem gets much smaller.

2. Vector indexes and chunk stores

RAG features embed customer documents into a vector database. The embedding plus the chunk text is customer data, and embeddings are not anonymous: the original text sits next to them, and research has shown text can be partially reconstructed from embeddings alone. Store tenant ID and source document ID as metadata on every vector so you can find them. If you already isolate tenants per namespace or per table, as described in our guide to multi-tenant RAG data isolation, a whole-tenant deletion becomes a single drop. Deleting one user's documents inside a shared tenant still needs document-level IDs.

One practical note for pgvector users: a DELETE marks rows dead, and the space and index entries are reclaimed by vacuum. The data is no longer returned by queries, but if your contract or policy talks about physical removal, make sure autovacuum is running and healthy on those tables.

3. Caches

Semantic caches, response caches and precomputed summaries all store model output that was generated from customer data. A cached summary of a customer's account is their data. Key caches by tenant, or at least include tenant ID in the cache key so you can purge by prefix. Short TTLs help here too.

4. Objects stored at your model provider

This is the copy teams most often miss, because it lives outside your infrastructure. Check provider documentation rather than assuming. Two current examples:

  • OpenAI says API abuse-monitoring logs are kept for up to 30 days unless the law requires longer, and API data is not used for training unless you opt in. Separately, stateful objects such as files, vector stores, assistants, threads, conversations, fine-tuning jobs and batches are kept until you delete them, and those endpoints are not covered by Zero Data Retention.
  • Anthropic says it deletes API inputs and outputs within 30 days by default, with exceptions that include services where you control retention, like the Files API, custom zero-retention agreements, and inputs flagged for usage policy violations.

The takeaway: the transient request data usually expires on its own within about 30 days. Anything you explicitly uploaded or stored through a provider API stays until your code deletes it. If your feature uses a hosted vector store or uploads files, your deletion job must call the provider's delete endpoints and record the response. Sources: OpenAI data controls and Anthropic data retention.

5. Eval datasets and fine-tuning data

Good teams turn real production failures into test cases, as we describe in building an eval set from production failures. That is the right practice, but it copies customer content into a repo or a dataset tool, often with no link back to the tenant. Fix this at the moment of capture: every eval example carries the source tenant ID, and you prefer synthetic or redacted versions of real cases. Fine-tuning is harder. Once a customer's data is in a trained model's weights, you cannot delete it from that model; you can only retrain without it. For that reason we avoid fine-tuning on raw customer data unless the contract explicitly allows it and the retraining cost is accepted up front.

6. Human review queues and analytics

Thumbs-down feedback, flagged outputs sent to a labeling tool, product analytics events that captured a prompt string. Each is a copy. Analytics events in particular often capture free text by accident through autocapture.

Design deletion as one tracked job, not a checklist

A wiki checklist that someone runs by hand will drift the first time a new store is added. Treat deletion as code.

Register every store in one place

Keep a small registry in code: each AI data store implements a deleteForTenant(tenantId) and, where needed, deleteForUser(tenantId, userId). When an engineer adds a new cache or a new provider-side object, code review asks whether it is registered. A store that cannot implement delete-by-key is a design problem to fix before launch, not after the first request.

Fan out, retry, and record

The deletion request creates one job record. The job calls every registered store, retries failures with backoff, and writes the outcome per store: deleted count, provider response ID where there is one, and timestamp. Stores with natural expiry, such as 30-day logs, record their expiry date instead of a delete call. When the customer or their auditor asks, you produce that record. This is the same discipline as an audit log for AI features, applied to deletion.

Verify, then test it in CI

After the job finishes, run a verification query against each store for the tenant ID and expect zero results. Then add an integration test that seeds a fake tenant across every store, runs deletion and asserts nothing remains. That test is what keeps the map honest as the feature grows.

Decide what backups mean

Backups usually cannot be edited record by record. The common approach is to let backups age out on a fixed schedule, make sure restored backups replay the deletion log before going live, and say this plainly in your data processing terms. Do not promise instant deletion from backups if your system cannot do it.

A launch checklist for any new AI feature

  • Every store the feature writes to is listed, with an owner and a deletion key.
  • Tenant ID is present on every log line, vector, cache key and eval example.
  • Provider-side objects (files, vector stores, threads) are tracked by ID in your database so you can delete them.
  • Log and cache retention periods are set and documented, ideally shorter than your contractual deletion window.
  • No raw customer data goes into fine-tuning without explicit contractual approval.
  • A deletion integration test runs in CI.

If you bring in an outside team to build the feature, the same map applies to their tooling and access. We cover that side in how to give an outside AI team data access without leaking customer data. When a deletion job fails in production, treat it like any other AI incident, as in our incident runbook for AI features.

FAQ

Do embeddings count as personal data?

Treat them as if they do. They are derived from the customer's text, they are usually stored next to that text, and they can leak information about it. Deleting the source document but leaving its vectors is not a complete deletion.

Does my LLM provider train on the data I send through the API?

For OpenAI and Anthropic commercial APIs, the published default is no, unless you opt in. That is separate from retention: both keep request data for a period for abuse monitoring, and anything you store through their APIs stays until you delete it. Check the current terms for every provider you use.

How fast do we need to complete a deletion request?

Under GDPR you must respond within one month of the request, with a possible extension for complex cases, and erase without undue delay when the right applies. Your customer contracts may set a different window. Design for the strictest one you have signed.

Can we remove one customer's data from a fine-tuned model?

Not reliably. Removing data from trained weights is still a research problem. The dependable option is to retrain without that data, which is why fine-tuning on raw customer data needs a clear decision before it happens.

Work with Boundev

Put expert judgment to work.

Talk to our team about evaluating AI-generated code, comparing responses, and building clear rubrics.