← Back to writing

Stop your AI feature from signing other people's names

Short answer: generative models will happily attach real people's names, signatures, quotes, and brand marks to content those people never made, because those patterns are common in their training data. To stop your AI feature doing it, treat any named person or brand in the output as an unverified claim: check it against the inputs, strip signatures and bylines by default, and label generated content as generated.

This week a widely shared Hacker News thread was about a consumer image model adding what looked like real cartoonists' signatures to imitation cartoons. It is a consumer product story, but the failure mode is the same one we see inside B2B SaaS features: an AI writing assistant that signs an email from the wrong executive, a review summarizer that quotes a customer who never said that, a marketing generator that puts a partner's logo on a slide. None of these are hallucinations in the usual RAG sense. The facts may be fine. The problem is who the output says it came from.

Why models produce false attribution

Attribution is a style feature as much as a fact. Cartoons end with a signature, op-eds end with a byline, testimonials end with a name and a job title, emails end with a sign-off. A model trained on millions of examples learns that a well-formed artifact of that type includes the attribution, and it fills the slot with something plausible. Often the most plausible thing is a real name strongly associated with the style.

Three conditions make it more likely in a product:

  • The prompt asks for a complete, finished artifact ("write a testimonial", "draw a New Yorker style cartoon", "draft a press quote from our CEO").
  • The prompt names a style that belongs to a person or brand ("in the style of", "like our partner's announcement").
  • Nothing downstream checks the output for names or marks that were not in the input.

The third one is the only one you fully control, so that is where the engineering goes.

Where this shows up in SaaS features

Writing assistants and email drafting

Sign-offs, signatures, and "on behalf of" lines get filled with whichever name appears most in the context window. In a shared inbox product, that can be a colleague, a previous sender in the thread, or the customer.

Review, feedback, and testimonial tools

Summaries of customer feedback often get turned into pull quotes. If the model paraphrases and then attributes the paraphrase to a named customer, you have published a quote that person did not say. In the US, the FTC's rule on consumer reviews and testimonials specifically targets fake or AI-generated reviews attributed to people, so this is not only a trust issue.

Image and slide generation

Logos, watermarks, and signatures are visual patterns. Image models reproduce them, sometimes recognizably. A generated slide with a real partner's logo, or an illustration with a real artist's signature, is the visual version of a fake quote.

Agents that send things

Once an agent can post, email, or publish, false attribution leaves your product and lands in someone's inbox. This is where the cost of one bad output is highest, and where containment matters most, as covered in the guardrails an AI feature needs before launch.

A practical defense in four checks

You do not need a research project. These checks are cheap, deterministic where possible, and fit into the output-filtering step most teams already have.

1. Grounded entity check

Run named-entity recognition on the output and on the inputs (prompt, retrieved documents, user profile). Any person or organization name in the output that does not appear in the inputs is flagged. An off-the-shelf NER model is enough to start; spaCy or a small hosted model handles person and organization names well in English. For a writing assistant, a flagged name can be highlighted for the user. For an autonomous send, it should block.

We use the same pattern for numbers and dates in RAG answers: the output may only contain specifics that can be traced to a source. Our post on reducing hallucinations in production RAG covers the grounding side.

2. Attribution slots are filled by code, not the model

Signatures, bylines, sign-offs, and quote attributions should come from your data, not the model's text. Ask the model for the body only, then append the signature from the authenticated user's profile. For pull quotes, require the model to return the source record ID with each quote, and render the name from that record. If the model cannot point to a record, the quote does not ship.

3. Verbatim check on quotes

Anything in quotation marks attributed to a person should match the source text, either exactly or above a high similarity threshold. Paraphrases are fine as long as they are not presented as quotes. This is a string comparison, not an LLM call, and it catches most fabricated testimonials.

4. Visual marks and style requests

For image features, block or rewrite prompts that ask for a named living artist's style or a named brand's mark unless the user owns it, and run a logo or text detector on outputs that might contain signatures. Sign generated images with C2PA content credentials so downstream viewers can see they were AI-generated.

Disclosure is becoming a requirement, not a nicety

Labeling generated content used to be a design choice. It is moving into law. The EU AI Act's transparency obligations in Article 50 began applying on 2 August 2026. Article 50(2) requires providers of systems that generate synthetic text, image, audio, or video to mark outputs in a machine-readable way, with a transition to 2 December 2026 for systems already on the market before August. US rules are more piecemeal, but the FTC rule above and state right-of-publicity laws already cover using a real person's name or likeness without consent. If you sell to EU customers or publish generated content on their behalf, plan for a visible label and machine-readable metadata.

The good news is that the same checks that make you compliant also make the feature more trustworthy, which is the point of a trust layer built from citations and confidence.

How to test it before launch

Add an attribution suite to your eval set. A useful starting set is 30 to 50 prompts:

  • Requests for finished artifacts that normally carry a name: testimonials, press quotes, signed letters, cartoons, op-eds.
  • Style requests naming real people and brands.
  • Threads with several people in context, to see which name the model picks for a sign-off.
  • Feedback summaries where the customer said something negative, to see whether the quote softens.

Pass criteria are deterministic: zero ungrounded person or brand names, zero quotes that fail the verbatim check, signatures always from the profile. Run it in CI next to your other evals, the same way we suggest for a red-team test suite. It is a small suite, and it catches the kind of mistake that ends up as a screenshot on social media.

Frequently asked questions

Is false attribution the same as hallucination?

It is a kind of hallucination, but it often happens when the content itself is fine. The model gets the facts right and the source wrong. Standard RAG grounding checks may not catch it unless they include names and quotes.

Can a system prompt fix this?

A system prompt telling the model not to invent names reduces the rate but does not remove it. Treat prompting as the first layer and the entity and verbatim checks as the layer you rely on.

Do I need this for an internal tool?

Less urgently, but yes if the tool drafts anything that leaves the company, such as emails, proposals, or social posts. Internal drafts become external content quickly.

What about generated images that only resemble a style?

Style alone is a legal gray area that varies by jurisdiction. Signatures, logos, and watermarks are not. Blocking those is a clear, cheap win; for style requests naming living artists, many teams choose to rewrite the prompt to a neutral description.

How long does this take to build?

For a text feature, the entity check, code-filled signatures, and verbatim quote check are a few days of work for one engineer, plus the eval suite. Image checks take longer because detection is fuzzier.

Get shipped

Rather we just build it?

Book a free scoping call and we'll ship your production-safe AI feature this week.