Your AI feature passed on demo data. Test it on production data
Answer first: an AI feature that passes on demo data has proven that the model can do the task, not that the feature works. Demo data is short, clean, English, well-formatted, and chosen by the person who wrote the prompt. Production data is long, truncated, pasted from email threads, half in another language, full of internal shorthand, and missing the fields the prompt assumed. Before you commit to a ship date, build a production-shaped test set: a few hundred real inputs sampled to match the actual distribution, redacted so they are safe to send to a model, and scored against what a good answer looks like. It usually takes two or three days and it is the single best predictor of whether the launch week is quiet or not.
Why demo data lies to you
Nobody builds a demo on the ugliest records in the database. The engineer picks a well-written support ticket, a tidy contract, or a complete customer profile, because the point of the demo is to show the idea working. The prompt gets tuned against those examples. Every edit makes it better on the demo set and tells you nothing about the rest.
Then the feature meets real input. Common surprises, in rough order of how often we see them:
- Length. The demo ticket was 200 words. Real tickets include a 40-message thread with quoted replies and signatures, and the context either overflows or the important sentence is buried on page six.
- Missing fields. The prompt said "the customer's plan is {plan}" and 30 percent of accounts have no plan recorded, so the model sees "the customer's plan is None" and reasons about it.
- Format drift. Dates in three formats, currency with and without symbols, HTML pasted into plain-text fields, CSV cells containing newlines.
- Language and tone. Partial Spanish, internal acronyms the model has never seen, sarcasm the classifier reads literally.
- Empty and near-empty inputs. A ticket whose body is "see attached," a note that is only a URL, a document that is one scanned image.
None of these are model failures. They are distribution failures, and the only way to find them before customers do is to test on data that has the same distribution as production.
How to build a production-shaped test set in two days
Day one: sample, do not cherry-pick
Pull a random sample from the real table the feature will read, not from a fixture file. Two to three hundred rows is enough for a first pass on most single-record features. Stratify it so the long tail is represented: by length bucket, by tenant size, by created date, and by whatever field most changes the input shape. If 5 percent of records are over 10k tokens, make sure 5 percent of the sample is too. Include the empties and the malformed ones on purpose. The sample should make you slightly uncomfortable to look at.
For a multi-tenant product, sample across tenants rather than taking one friendly customer's data. Tenants differ in how they use free-text fields, and the feature has to work for the ones who use them badly. The tenant boundary also has to hold during testing, which is covered in the data isolation rules for multi-tenant RAG in SaaS.
Day one, afternoon: make it safe to send
Real data contains names, emails, phone numbers, and sometimes payment or health details. Before any of it goes to a model provider, run it through the same redaction step the production feature will use. If the production feature does not have one yet, this is when you build it, because the test set is the first real traffic. The PII redaction gateway pattern describes a placement that works for both testing and production, so you do not end up with two different redactors that disagree.
Check your data processing terms with the customer contracts before using live records, even redacted. Some enterprise agreements restrict secondary use. If a subset of tenants is off limits, exclude them at sampling time and note it in the test set's README.
Day two: define what good looks like, then score
A test set without a scoring rule is just a pile of inputs. For each sampled row, decide how you will judge the output. Three practical options, in increasing cost:
- A rubric with pass and fail criteria that a reviewer applies by hand. Cheapest to set up, slowest to run, fine for a first pass of a few hundred rows.
- Reference answers written by a domain expert for a subset, with an automated similarity or exact-match check for structured outputs like tags and scores.
- A model-graded check with the rubric in the prompt, validated against the hand-graded subset so you know its agreement rate before trusting it.
Run the feature over the whole set and look at the failures by bucket, not the aggregate. An 88 percent pass rate is meaningless if it is 99 percent on short inputs and 40 percent on long ones, because the long ones are the tickets that matter most. The bucketed failure table is what you take into the planning meeting, and it turns into the acceptance criteria and evals for the feature that gate the release.
What the results change
Usually the test set changes the design, not just the prompt. The long-input bucket forces a decision about truncation, chunking, or summarizing before the main call. The missing-field bucket forces a decision about whether to skip the feature for incomplete records or to prompt without the field. The empty-input bucket forces a decision about what the UI shows when there is nothing to work with, which is a product question nobody asked during the demo.
These decisions take hours when made before launch and days when made after, because after launch they come in as customer complaints, each one needing investigation to find the bucket it belongs to. The set also becomes your regression suite. Every prompt change, model upgrade, or retrieval tweak gets run against it, and the production failures you collect after launch get added to it, so the set grows toward the real distribution over time.
When a team sends us an AI feature to build, this test set is one of the first deliverables, before the feature is considered scoped. It is how we can quote a ship date with a straight face. The way that fits into the week-long task cycle is described in how a Boundev task goes from brief to production.
How many rows do I need in a production-shaped test set?
Two to three hundred stratified rows will surface the main failure buckets for a single-record feature. Go larger only when the feature has many distinct input types or when a bucket you care about has fewer than 20 examples, because below that the pass rate is noise.
Can I use synthetic data instead of real records?
Synthetic data is useful for filling a bucket that is rare in production or for cases you cannot legally sample. It is a poor substitute for the main set, because a model generating "messy" data produces a tidy imitation of mess. Use real data for the distribution and synthetic data for the gaps.
What pass rate is good enough to ship?
There is no universal number. Decide per bucket, based on what a wrong answer costs. A suggested tag the user can change can ship at 80 percent. A summary that goes to a customer without review needs to be far higher, or needs the review step added. The point of the test set is to make that a per-bucket decision with evidence rather than a gut call on launch day.
Rather we just build it?
Book a free scoping call and we'll ship your production-safe AI feature this week.