Giving an outside AI team data access without leaking customer data
Short answer: an outside AI engineering team rarely needs raw production data to build a good feature. Give them a staging environment loaded with masked or synthetic data, their own scoped credentials that you own and can revoke, and a signed data processing agreement before any real record leaves your systems. Reserve live data access for a narrow, logged, time-boxed window, usually during evaluation and launch.
Most founders hit this question the week they bring in outside help. The AI feature needs real examples to work well, the engineers are asking for "a database dump to get started", and nobody has decided what is allowed. The default answer in a hurry is to share an admin login. This post is the setup we would want if we were the customer, written from the side that usually receives the access.
Why AI work raises the stakes on data access
A normal outsourced feature touches your data through code. An AI feature touches it in three more places, and each one is a place customer data can end up that you did not plan for.
Prompts and model providers
Anything placed in a prompt is sent to a model provider. If the external team runs experiments from their own provider account, your customer records now sit in a third party's logs under an account you do not control. The major API providers state that API inputs are not used for training by default (see OpenAI's enterprise privacy page), but retention and abuse-monitoring logs still exist, and the contract is between the provider and whoever owns the key.
Eval sets and fixtures
Good AI work produces an evaluation dataset: a few hundred real inputs with expected outputs. That file is valuable and it is often copied into a repo, a notebook, or a spreadsheet. If it was built from production rows, it is a copy of production data that will outlive the engagement.
Vector stores and caches
A retrieval feature embeds your documents into a vector index. Embeddings are not anonymization; the source text is usually stored next to the vector so it can be shown to the model. A test index built on a laptop is a second copy of your knowledge base.
The setup that works
The pattern below is boring on purpose. It takes a day or two to put in place and removes most of the risk before any code is written.
1. Paperwork before access
Sign a data processing agreement before the first record moves. If you have EU or UK users, GDPR Article 28 requires a written contract with any processor handling personal data on your behalf. If you handle health data in the US, the external team needs a business associate agreement before touching PHI. The same agreement should say who owns the prompts, eval sets, and code; our post on who owns your AI code when an engagement ends covers that half.
2. Accounts and keys you own
Every credential the team uses should live in your organization: your cloud account, your model provider organization, your secrets manager. Create a separate provider project for the engagement so usage, spend, and logs are visible to you and can be shut off in one action. Never let an outside team run your customers' data through their own API keys. This is the same least-privilege idea we apply to agents themselves in scoped credentials for production AI agents.
3. A staging environment with masked or synthetic data
For most of the build, the team should work against a staging copy where personal fields are masked or replaced. Open-source tools such as Microsoft Presidio detect and redact names, emails, phone numbers, and similar fields in free text, which matters because AI features usually read free text: support tickets, notes, chat logs. Pair it with simple column-level masking in SQL for structured fields. The aim is data that keeps the shape and messiness of production without the identities, which is the same idea behind production-shaped test data.
4. Narrow, read-only live access when it is genuinely needed
There are moments where masked data is not enough: checking retrieval quality on the real corpus, or measuring accuracy before launch. For those, grant read-only access to a replica or a database view that exposes only the columns the feature uses. Put an end date on it. Log the queries. If you use Postgres, row-level security and a dedicated role per person make this straightforward and auditable.
5. Eval data lives in your systems
Decide up front where the evaluation dataset is stored: a bucket or repo you own, with access granted to the team, not the other way around. If it was built from real records, treat it as production data and apply the same retention rules. When the engagement ends, the eval set stays with you and the access goes away.
6. An offboarding step written on day one
Write the exit list at kickoff: revoke provider keys, remove cloud roles, delete laptop copies of eval files and local vector indexes, and confirm in writing. It takes ten minutes to write when nobody is leaving and is easy to forget when someone is.
What to give, by phase
A simple way to keep this proportionate is to tie access to the phase of the work.
| Phase | Data the team needs | Access level |
|---|---|---|
| Scoping and prototype | 20 to 50 hand-picked, masked examples | Shared folder you own |
| Build | Masked or synthetic staging data at realistic volume | Staging environment, scoped keys |
| Evaluation | A few hundred real inputs, ideally de-identified | Read-only view, time-boxed, logged |
| Launch and monitoring | Production traces with personal fields redacted | Observability tool access, no direct DB writes |
If a team asks for full production access in the first week, ask what specific question they are trying to answer. There is usually a smaller dataset that answers it.
Common mistakes we see
- Sharing one admin login across the whole outside team, which makes every action unattributable and every revocation a password change for everyone.
- Running experiments through the vendor's model provider account, so your data and your spend are in someone else's logs.
- Treating embeddings as anonymized data and copying a production vector index to a laptop.
- Building the eval set in a personal spreadsheet that nobody remembers to delete.
- Waiting until the end of the engagement to think about offboarding.
None of these come from bad intent. They come from moving fast without a default. Setting the default is the client's job, and a good partner will be relieved you did.
How this fits a subscription model
On a subscription engagement like ours, engineers work in your repos and your cloud accounts from the start, which makes the setup above the normal path rather than an extra step. You can see how tasks move from request to shipped code on how Boundev works, and the reliability and security commitments are on our reliability page. If you are comparing models, dedicated AI engineering teams follow the same access rules.
Frequently asked questions
Does an outside AI team ever need real production data?
Sometimes, mainly to measure quality before launch. Retrieval and classification features can behave differently on the real corpus. Grant that access read-only, scoped to the needed columns, time-boxed, and logged, and keep the resulting eval set in your own systems.
Is it safe to send customer data to an LLM API?
It can be, if the API account belongs to you, your provider terms cover your data type, and you have checked the provider's retention settings. The risk is usually not the provider; it is data flowing through accounts and copies you do not control.
Are embeddings anonymous?
No. Vector stores typically keep the original text alongside each vector so it can be passed to the model, and research has shown text can be partially reconstructed from embeddings alone. Treat a vector index with the same care as the documents it was built from.
What should the contract say about data?
At minimum: the processing purpose, the data categories, security measures, subprocessors, breach notification, and deletion or return at the end. For EU or UK personal data that is required by GDPR Article 28. Add ownership of prompts, eval sets, and model artifacts.
How long does this setup take?
For a typical SaaS product, a day or two: a provider project, a few scoped roles, a masked staging copy, and a signed agreement. That is small next to the cost of cleaning up a copy of production data that ended up somewhere it should not be.
Rather we just build it?
Book a free scoping call and we'll ship your production-safe AI feature this week.