The real hallucination numbers
Somewhere today, a prospective customer asked an AI assistant about your company. The assistant answered fluently, confidently, and maybe partly wrong. It might have described a product you discontinued, quoted a price you never charged, invented an integration you do not have, or attached someone else's complaint to your name.
The first version of this article leaned on an unverified claim about hallucination rates. The better lesson is more useful: hallucination depends on grounding. Vectara's Hallucination Leaderboard evaluates whether a model stays faithful while summarizing a provided source document. That is the good lane. The model has the document in hand, the answer can be checked against it, and the best systems now keep error rates low.
Brand search is not always that lane. A buyer does not ask with your latest pricing page attached. They ask an open-ended question into ChatGPT, Gemini, Claude, Perplexity, or an AI answer box. The system retrieves what it can, weighs whatever the public corpus says, and fills gaps when the corpus is stale, thin, or contradictory. Your risk lives in the distance between a grounded summary and an answer assembled from weak brand evidence.
So the practical question is not whether LLMs hallucinate in the abstract. It is whether the answer engine has a current, clear, machine-readable source for every material claim about your business. When it does, the model has something to summarize. When it does not, it completes the pattern. Fluently. About you.
The case that settled owned-surface liability
Start with the owned surface because the precedent is already plain enough for a boardroom. In Moffatt v. Air Canada, a traveler asked the airline's website chatbot about bereavement fares. The bot gave the wrong answer and told him he could apply retroactively. The real policy said the opposite.
Air Canada tried to separate itself from the chatbot's statement. The tribunal did not let that argument carry the day. The chatbot was part of the company's website, and the company was responsible for the information provided there. The amount was small. The operating principle was not.
A support bot, product finder, AI concierge, or policy helper is a promise-generating surface. If it improvises on refunds, compatibility, eligibility, safety, pricing, or regulated claims, it can publish language no human approved. That makes hallucination governance a marketing, legal, and operations problem at the same time.
The control is not to forbid AI support. The control is to ground the bot in a verified corpus, make refusal safer than invention, log the answer and source, and escalate when the question crosses policy boundaries. If the system cannot show where the answer came from, it should not answer as the company.

Inbound hallucinations are harder
The harder problem runs on systems you do not deploy. A buyer asks an answer engine whether your product is reliable, how it compares with a competitor, what your pricing is, or whether your company has a controversy. You are not in that room, and there may be no click for your analytics to record.
The recurring failure modes are predictable enough to audit. Zombie facts keep old pages alive. Invented specifics add plausible capabilities. False associations blend your entity with a similarly named company or a category scandal. Averaged opinions turn fragments of review chatter into a verdict no source quite said.
Zombie facts
Old pricing, discontinued products, former executives, or superseded policies become the current answer.
Invented specifics
The engine creates plausible integrations, guarantees, regions, ingredients, certifications, or performance claims.
False associations
A similarly named company, competitor incident, category controversy, or unrelated lawsuit gets blended into your entity.
Averaged opinions
Reviews, forum posts, and summaries become a verdict no individual source quite said.
The Munich AI Overviews ruling matters because it gives the worst inbound falsehoods a legal address. As The Decoder reported, a German court found that Google could be responsible for a false AI Overview claim about a business. Most brand errors should be fixed before lawyers enter the room, but the escalation path changes the seriousness of the file you keep.
Why brands are structurally exposed
Brand hallucination is rarely random. It is often the machine papering over a gap the company left in public. Thin product pages, undated policy pages, contradictory pricing claims, soft launch copy, old PDFs, orphaned help articles, and weak entity signals all create space for the model to infer.
This is why GEO and AEO work are not only visibility plays. They are defensive infrastructure. The true version of every material fact about the brand should be easy to find, parse, date, quote, and reconcile across owned and high-authority third-party sources.
A good source corpus is plain before it is poetic. It states the current product, supported regions, pricing posture, integrations, policies, safety constraints, compliance status, and discontinuations in language a machine can lift without guessing. It uses consistent entity names and schema. It does not hide critical facts in graphics, PDFs, or vague campaign slogans.
The uncomfortable truth is that many brand sites are designed for persuasion, not retrieval. They give the model adjectives when it needs nouns, screenshots when it needs facts, and narrative when it needs a dated source of record. Then the model invents the missing connective tissue.
The audit: measure what machines say
You cannot correct what you have not catalogued. Start with a standing prompt panel: 20 to 40 questions a real buyer, journalist, partner, candidate, or regulator might ask about the company. Include pricing, product fit, comparisons, complaints, policies, locations, safety, litigation, and the five topics you would least want an assistant to improvise.
Run that panel across the engines your audience actually uses. For most brands, that means ChatGPT, Gemini, Claude, Perplexity, and Google AI Overviews at minimum. Log answers verbatim with dates, engine, prompt, source links, and screenshots. A visibility tool can automate parts of this. A disciplined spreadsheet works on day one.
Triage by severity, not volume. Tier one errors create legal or safety risk. Tier two errors affect revenue: pricing, availability, capabilities, comparisons. Tier three errors are staleness, framing, and tone. Track an accuracy score by engine over time. It is the reputational twin of share of answer.
The correction playbook
First, fix the corpus. For every recurring error, publish or repair the authoritative page that should have grounded the answer. Make it plain-language, dated, structured, and entity-consistent. The answer engines cannot cite the source you never gave them.
Second, use platform correction channels. AI Overview feedback, model-provider report mechanisms, and business or publisher channels are imperfect, but they establish notice. Keep screenshots, dates, prompts, false outputs, corrected source URLs, and follow-up attempts.
Third, cage your owned bots. Use retrieval-grounded responses from a verified policy corpus. Prefer refusal over invention on policy, pricing, safety, eligibility, and regulated claims. Log the sources shown to the model and the answer sent to the user. Escalate to a human when the bot lacks evidence.
Fourth, reserve the legal ladder for persistent tier-one errors. Most brands will not need it. The point of the file is to make escalation credible when an engine keeps publishing a damaging falsehood after correction attempts.
Report the work like a KPI: accuracy by engine, open tier-one items, corrections shipped, time to correction, and source pages strengthened. The moment leadership sees the machine-told version beside the campaign-told version, this stops being an oddity and becomes maintenance.

