How to red-team your AI feature in CI before attackers do
Short answer: red-teaming an AI feature means attacking it on purpose, with prompt injection, data-leak attempts, tool abuse and cost bombs, before your users or a security reviewer do. For a SaaS team, the version that lasts is not a one-off pentest. It is a few hundred attack cases kept in the repo and run in CI on every prompt, model or tool change, with a hard fail on any regression.
Agent platforms are shipping red-teaming as a standard feature this fall, and enterprise security questionnaires now ask about it by name. Here is how we set up a red-team suite for a production AI feature: what to attack, how to grade it, and how to keep it from rotting.
Why a one-time red-team exercise is not enough
A classic pentest checks a system that changes slowly. An AI feature changes every time someone edits a prompt, swaps a model version, adds a tool or widens retrieval to a new data source. Each of those can reopen a hole that was closed last month. Some examples of how this happens in practice:
- A teammate shortens the system prompt to cut tokens and deletes the line that told the model never to reveal other tenants' data.
- A new model version follows instructions inside retrieved documents more eagerly than the old one did.
- A "send email" tool is added to an assistant that already reads untrusted inbound email, and nobody connects the two.
None of these show up in unit tests, and none of them show up in a quality eval that only asks "was the answer helpful". You need tests written from the attacker's side, and they need to run as often as the feature changes. If you already run quality evals in CI, as described in our guide to catching regressions on model updates, the red-team suite is a second test set on the same pipeline.
What to attack: a SaaS-shaped threat list
The OWASP Top 10 for LLM Applications (2025) is the standard checklist, and security reviewers will expect you to map to it. For a typical B2B SaaS feature, five categories cover most of the real risk.
Prompt injection, direct and indirect (LLM01)
Direct injection is a user typing "ignore your instructions". Indirect injection is worse for SaaS: the attack sits inside content your feature reads, such as a support ticket, a PDF upload, a CRM note or a web page. Write cases where the malicious instruction lives in the retrieved or uploaded content, not the chat box. We go deeper on defenses in our prompt injection post.
Cross-tenant and sensitive data disclosure (LLM02)
This is the one that ends enterprise deals. Seed your test environment with two tenants, plant a unique canary string in tenant B's data (for example CANARY-7f3a91), and run attacks as tenant A that try to pull it out. Any response containing the canary is an automatic fail, with no judgment call needed.
System prompt leakage (LLM07)
Assume your system prompt will leak eventually, and test that leaking it does no damage. Put a canary in the system prompt too. If extraction attempts succeed, the real check is whether the prompt contained anything that matters, like API keys, internal URLs or business rules you would not show a customer. It should not.
Excessive agency (LLM06)
For every tool the model can call, write cases that try to make it call that tool when it should not. Get it to send an email to an address from the document, delete a record, or change a setting for another user. Grade on the tool calls in the trace, not on the text reply. A model that politely says "I can't do that" while issuing the call has still failed.
Unbounded consumption (LLM10)
Inputs designed to make the model generate huge outputs, loop through tools, or pull an entire knowledge base into context. These are cost attacks as much as availability attacks. Pass criteria are numeric: the response stays under a token budget, the agent stops within N steps. Per-tenant caps, covered in our multi-tenant cost caps post, are the backstop, but the suite should prove the feature does not hit them under normal attack inputs.
Building the suite
Seed cases by hand, then generate variants
Start with 20 to 40 hand-written attacks that are specific to your product: your tool names, your document types, your tenant model. Generic jailbreak lists miss the attacks that matter for your feature. Then use an attacker model to generate variants of each seed (paraphrases, other languages, encodings, attacks split across two messages) to reach a few hundred cases. Open-source tools like promptfoo, garak and PyRIT can generate and run these, and most agent platforms now offer something similar. The tool matters less than having your own product-specific seeds.
Grade deterministically wherever you can
An LLM judge that decides whether a response "seems unsafe" is noisy, and noisy tests get ignored. Prefer checks that cannot be argued with:
- A canary string from another tenant or from the system prompt appears in the output.
- A forbidden tool was called, or an allowed tool was called with an argument outside an allowlist.
- Output tokens or agent steps exceeded a fixed budget.
- The output contains markup your frontend would render unsafely, such as a script tag or a markdown image pointing to an outside domain (a common exfiltration trick, and an Improper Output Handling issue under LLM05).
Keep an LLM judge only for the leftovers, such as "did the assistant give harmful advice", and calibrate it against 50 or so hand-labeled cases before trusting it.
Run it against the real wiring
Red-team the feature as it runs in production: the same retrieval, the same tool definitions, the same output rendering. Use a staging environment with fake tenants and stub tools that record calls instead of acting on them. Testing the bare model with your system prompt misses most indirect injection and tool abuse, because those live in the wiring.
Wiring it into CI
A full run of a few hundred cases against a frontier model takes minutes and costs a few dollars. That is cheap enough to run on every pull request that touches prompts, tool definitions, retrieval config or the model version. Split it into two tiers:
- A fast smoke set of about 50 high-severity cases on every relevant pull request, which blocks the merge on any failure.
- The full suite nightly and before each release, posting a pass-rate trend so slow erosion is visible.
Set the bar as zero failures on the deterministic checks for cross-tenant leakage and forbidden tool calls. For fuzzier categories, track the pass rate and fail the build if it drops by more than a couple of points from the last release. Model outputs vary, so run flaky cases three times and count a single failure as a failure. For security tests, one leak in three attempts is a leak.
Keep feeding it
Every incident, bug bounty report or odd production trace that looks like an attack becomes a new case, the same way production failures feed a quality eval set (see building an eval set from production failures). A suite that has not grown in three months is probably no longer testing the feature you ship.
What this does not replace
A red-team suite tells you whether known attack shapes get through. It does not replace the controls that make attacks harmless when they do get through: least-privilege tools, human confirmation on destructive actions, tenant isolation enforced in the data layer instead of the prompt, and output sanitizing. Start with those, which we cover in our pre-launch guardrails checklist, then use the suite to prove they still hold after every change.
FAQ
How many red-team cases do we need before launch?
There is no magic number. A useful starting point is 20 to 40 product-specific seeds that cover each tool and each untrusted input source, expanded to a few hundred variants. Coverage of your actual tools and data paths matters more than raw count.
Should we hire an outside firm instead?
Both help. An outside red team is good at finding attack classes you did not think of. The in-repo suite makes sure those findings stay fixed. Turn every outside finding into a CI case.
Will the suite slow down our releases?
The pull-request smoke set adds a few minutes, and only on changes that touch AI behavior. Teams usually notice the opposite: changing prompts gets less scary, so they ship prompt changes more often.
Who should own it?
The engineers who own the feature, not a separate security team. They know which tools and data paths changed. If your team is stretched, our LLM engineers can set up the first suite and the CI gates alongside your existing evals.
Rather we just build it?
Book a free scoping call and we'll ship your production-safe AI feature this week.