Skip to content

Guide

Auditing AI Accuracy: How to Check What ChatGPT, Gemini and Perplexity Really Say About Your Brand

35% of chatbot answers now repeat a false claim, up from 18% a year earlier. A complete method to audit what generative AI says about your brand, and why manual audits hit their limit fast.

3 August 2026
9 min
Nanga Team

For two years, the question brands asked was: "does ChatGPT know me?". That is changing. The real question is becoming: "does ChatGPT say accurate things about me?".

The distinction matters. Being absent from an answer costs an opportunity. Being misrepresented costs a sale, sometimes a reputation.


The problem is no longer absence, it is misrepresentation

A recent NewsGuard audit, which tests roughly ten demonstrably false claims on the leading chatbots every month, found that the share of answers repeating a false claim rose from 18% to 35% between July 2024 and August 2025. Nearly double in a year, even as the models improve on other fronts and now answer close to 100% of questions instead of cautiously abstaining as they did in 2024.

Within that same audit, the gaps between models are significant.

ModelShare of answers repeating a false claim
Pi (Inflection)> 45%
Perplexity> 45%
ChatGPT~ 40%
Meta AI~ 40%
Gemini< 17%
Claude< 17%

Source: NewsGuard, monthly AI chatbot audit, August 2025.

A more recent audit, run in French in April 2026 on Mistral's Le Chat, goes further: when a user simply asks for clarification on a rumor, without trying to manipulate the model, the rate of false-information repetition climbs to 70% in French versus 60% in English.

In other words, the guardrails are weaker on French-language content. A point worth keeping in mind for any brand tracking its presence across AI systems in non-English markets.

These audits cover news and sensitive topics. But the same mechanism applies exactly the same way when a model answers a question about a company: its products, its pricing, its service area, its recommendations. That is where accuracy stops being a theoretical concern and becomes a business one. All the more so given that half of all ChatGPT conversations are decision-related, as the NBER data on real usage shows.


Why AI gets brands wrong in particular

The cited sources are not yours

An LLM does not "know" anything by heart about a given company. When it answers a brand question, it most often relies on retrieval-augmented generation (RAG), pulling passages it deems relevant before drafting an answer.

An Omniscient Digital analysis covering more than 23,000 LLM citations shows that brand-owned websites do not dominate those citations. Media outlets, forums and third-party review platforms do, accounting for close to half of the sources retained on brand-related queries.

In practice: a well-maintained corporate site carries less weight than what is said elsewhere on the web.

Models are rewarded for guessing

On top of that sits a bias documented by OpenAI itself in a research report: standard evaluation procedures reward an attempted answer, even a wrong one, over an admission of uncertainty. That structurally pushes models to guess rather than say they do not know.

An outdated price, a discontinued product line, a competitor named instead of you: these are not isolated bugs, but a predictable consequence of how the models work.


What a brand accuracy audit actually covers

Unlike a NewsGuard-style audit centered on news misinformation, an accuracy audit applied to a brand answers more operational questions:

  • Factual accuracy: does the model correctly describe your products, your pricing, your service area, your positioning?
  • Source attribution: when the model cites a source, is it your own content or a third party, often older or approximate?
  • Competitive confusion: does the model confuse you with a competitor, or recommend one in your place on your own queries?
  • Stability over time and across engines: is the answer the same on ChatGPT, Gemini, Perplexity and Claude, and does it hold from one week to the next?

These four dimensions map to what Recommendation Visibility measures: how often, and how well, a brand appears when a user expresses intent inside an AI system. That is the focus of the Search & AI Visibility pillar.


How to run the audit, step by step

Several published methodologies converge on the same outline, whether run by hand or with a tool.

1. Build a representative query set

Brand, products, geographies, comparisons ("X vs Y"), recommendation queries ("what would you recommend for...") and category queries. Not just the company name: most recommendation opportunities play out on questions where the user never mentions the brand.

2. Query several engines, several times

Answers vary from one model to another and from one session to the next. A specialized agency recommends repeating every query two to three times before drawing a conclusion. A single answer, on a single engine, is not a measurement.

3. Qualify each answer at three levels

One published method distinguishes whether the brand is simply visible (it exists somewhere in the answer), mentioned (named explicitly) or recommended (highlighted with a link). At each level, factual accuracy is a separate question: a brand can be recommended on the strength of a false argument.

4. Repeat over time

A one-off audit gives a snapshot at a point in time. It hits its limit as soon as the goal is to track evolution or compare several competitors over months.


The limit of the manual audit

Step 4 is precisely what breaks the manual method.

Asking five questions by hand on ChatGPT gives a useful signal on a given day. But a serious accuracy audit means crossing dozens of queries, across several engines, repeated at regular intervals, with comparable history to detect degradation.

That is what the NewsGuard audits themselves confirm, run monthly since July 2024, precisely because an isolated measurement says nothing about the trend.

It is orchestration and tracking work, not writing work. That is where a monitoring tool becomes relevant.


Industrializing the audit with Nanga

This is what the Nanga GEO module covers. Rather than asking questions by hand once in a while, the platform continuously orchestrates a library of more than 500 prompts across 8 LLMs (ChatGPT, Claude, Gemini, Perplexity, Copilot, Mistral and others), captures the full history of the answers obtained, and automatically compares your brand to your competitors on every query.

In practice, that turns the four manual steps above into permanent tracking:

  • Visibility tracking per LLM and per query, with alerts as soon as a position change is detected.
  • Answer history, which surfaces when a model has started repeating incorrect information about your brand, instead of discovering it by chance.
  • Semantic analysis of competitor citations, to understand why a competitor is named in your place and which sources it relies on.
  • Dashboards that translate those gaps into marketing decisions rather than another score.

The AI Visibility Index aggregates these measurements into a unified score of presence across the AI ecosystem, and the Decision Engine turns the detected gaps into action priorities. That is the principle of the Marketing Intelligence System: connecting visibility in AI answers to investment decisions, rather than stacking indicators.

The goal is not to replace human vigilance, but to give it the regularity and scale a one-off audit cannot offer, on a subject where, as the NewsGuard figures show, the situation moves fast and not always in the right direction.


Where to start

Two sector studies show what this measurement looks like applied to a real market, with the prompts, the scores and the competitive gaps:

For an audit on your own brand and competitive set, request a demo.


Sources:

  • NewsGuard, The rate of false claims repeated by AI chatbots has nearly doubled in a year.
  • Silicon.fr, AI chatbots: information reliability is degrading.
  • Archimag, 35%: the rate of fake news relayed by AI chatbots has nearly doubled in a year.
  • Siècle Digital, Mistral AI faces a serious misinformation problem with its consumer chatbot.
  • Semjuice, AI hallucinations: protecting your brand from LLM errors.
  • Genn, ChatGPT visibility: 5 tests to audit your brand.
  • Agence MYA, AI visibility audit: test your presence in ChatGPT and Perplexity.
  • Olenx, Auditing your AI visibility for free.
  • toonetcreation, AI visibility audit: how to know whether your company is visible in ChatGPT, Gemini, Perplexity and Google AI Overviews.

See Nanga in action

Discover how the Marketing Intelligence System turns data into decisions.