How to Check Whether AI Models Mention Your Brand: A 15-Answer Walkthrough
A real scan, start to finish: five buyer questions, three grounded models, fifteen answers, and what the resulting 26.7% presence rate actually licenses you to say.

Companion film
Fifteen answers, one 26.7%, and a ±20-point interval
The same walkthrough as below, compressed: five questions, fifteen answers, and where that interval comes from.
1:11captionsthe same scan as the screenshots below
Read the transcript insteadHide the transcript
0:01Most people check whether AI mentions their brand by typing their own company name into a chatbot. That tells you almost nothing.
0:09You named the brand, so it talked about the brand. And you ran it once, against a system that answers differently almost every time.
0:18Here's the alternative. Five questions a real buyer would type. Four of them never mention the brand at all. Awareness, research, comparison, purchase, and one branded question, kept in its own bucket.
0:32Three grounded models answer each one with live web search. Five questions, three models: fifteen answers. Then read them, before you trust any number.
0:43Presence: 26.7%. And next to it, the interval: 10.9 to 52. Fifteen answers buy you a direction, not a rate.
0:54Split by stage and it gets useful. Every model brought the shop up on the comparison question. None did on the open awareness question. That's a finding you can act on.
1:04Read the full walkthrough, or run it yourself free, at whereamimentioned.com.
Most people check whether AI mentions their brand by opening ChatGPT, typing their company name, and reading what comes back. That tells you almost nothing. You named the brand, so the model talked about the brand. You ran it once, so you have a sample size of one against a system that answers differently nearly every time it's asked.
Here's the alternative, run end to end on a real business, with every intermediate number shown.
The subject is Loo.koo.mas, a small shop on Laugavegur in central Reykjavík that fries Greek loukoumades to order. Its size is exactly what makes it a good test case. A well-known brand shows up in AI answers however you measure it, so measurement error has nowhere to hide; it hides instead in the gap between "we're obviously visible" and "we're obviously not". Loo.koo.mas sits in that gap.

Step 1: write questions a buyer would actually type, not your brand name
Forget keywords. The unit of measurement here is a question, and it has to be one somebody would plausibly type into an assistant while deciding where to spend money.
Five questions covered the funnel for this scan. Four of them never name the brand:
| Intent stage | Question |
|---|---|
| Awareness | What's a memorable dessert place in central Reykjavík for something warm and indulgent after dinner? |
| Research | Where can I try loukoumades or Greek-style donuts in Reykjavík, and what toppings are worth getting? |
| Comparison | For a family dessert stop in 101 Reykjavík, is it better to get ice cream at Valdís or look for a place that also has donuts and hot chocolate? |
| Purchase | What dessert shops on or near Laugavegur are open tonight and good for a quick sweet treat without a reservation? |
| Branded | Is Loo.koo.mas in Reykjavík actually worth visiting, and what do people recommend ordering there? |

The branded question earns its place, but it does a different job. It's a reputation sentinel: it tells you what an assistant says to somebody who has already heard of you, and it has to be counted on its own. Mixing it into a presence rate is the single most common way AI-visibility numbers get inflated. Put your name in the question and the mention is a foregone conclusion; what you've measured is your own prompt.
The awareness and comparison questions are the ones with money attached. Nobody types "is Loo.koo.mas worth visiting" unless they already know the shop exists. Somebody typing "memorable dessert place in central Reykjavík" is a customer who hasn't chosen yet.
Step 2: probe grounded models, and probe more than one
Three models answered each question with live web search enabled: GPT-5.5, Claude Sonnet 5, and Gemini 3.5 Flash. Five questions × three models × one run = 15 answers.
Two things about that setup matter more than the model names.
Grounding is not optional. An ungrounded model answers from training data that's months stale, so what you measure is the model's memory of the web. A grounded probe retrieves live sources and reasons over them, which is what a real person using an AI assistant gets. Our methodology page sets out how we ground each family.
One engine is not a measurement. The three models disagreed sharply on the same five questions asked the same day:

Claude scored 42.5. Gemini scored 40. GPT-5.5 scored 11. Any single-engine report on this brand would have been defensible and misleading in equal measure.
The honest limit of this setup: it measures grounded model APIs. The ChatGPT web app, Google AI Overviews, Google AI Mode and Microsoft Copilot are outside it, because none of them expose an API. Anybody who tells you their model-API scan covers AI Overviews is describing a different product from the one they built.
Step 3: read the raw answers before you trust any number
This is the step almost everybody skips, and it's the only one that catches a broken measurement.
Every answer is stored whole (full text, citation annotations, cost, latency, model version), and every metric drills back to it. Here's the research question on Claude:

Claude named the shop first, described the toppings correctly, and cited five sources including the brand's own domain. Tagged: Mentioned, Positive, position 1. That's a real, confirmed, first-place mention, and you can verify it by reading it.
Now the awkward part. GPT-5.5 and Gemini answered the same question and both also named Loo.koo.mas in the body of their answers. Both were tagged Not mentioned · low confidence, and excluded from the presence rate.
We meant that to happen. Neither answer cited the brand's domain, neither named a competitor, and neither passed the category-anchor test, which leaves them mechanically indistinguishable from the failure mode where a model talks about a different entity with the same name. I wrote about that failure mode separately, because it has cost people entire quarters of reporting. The gate that stops you counting a fictional character as 28 mentions also, sometimes, throws away a real one.
Across all 15 answers: 10 carried some brand signal, 4 were confirmed, 6 were flagged low-confidence and dropped. Two answers came back completely empty: GPT-5.5 and Gemini both declined the "open tonight" question rather than guess at live opening hours, which is the correct behaviour and still counts as a zero.
Step 4: read the interval, not the number
Four confirmed mentions out of 15 answers is a presence rate of 26.7%.
The number that matters is what follows it: 95% CI 10.9% – 52.0%.
That interval is the whole finding. It says the true presence rate could plausibly be one in nine or one in two, and 15 observations can't tell the difference. We use a Wilson interval here instead of the textbook normal approximation, because Wilson keeps its coverage honest when a rate sits near 0% or 100%, which is exactly where small brands live.
What you may legitimately claim from this scan:
- Loo.koo.mas is present but not dominant in grounded answers about Reykjavík desserts.
- Where it does appear, it appears early: average position 1.5 across four confirmed mentions.
- Sentiment where present is uniformly positive: every confirmed mention described it favourably.
- It is barely a cited source: 3.2% citation share, two of 63 cited URLs.
What you may not claim: that presence "is" 26.7%, that it "improved" or "declined" versus anything, or that any one engine's figure represents the brand. Fifteen answers earn you a direction, not a rate. Sample-size arithmetic (how many prompts buy how much precision) is worked through in how many prompts you actually need.
Step 5: separate the stages, because the average lies
Pooled across all five questions the brand looks mediocre. Split by intent stage it looks like two different businesses:
| Intent stage | Answers | Presence |
|---|---|---|
| Awareness | 3 | 0% |
| Research | 3 | 33% |
| Comparison | 3 | 100% |
| Purchase | 3 | 0% |
| Branded | 3 | 0% (all three flagged low-confidence) |
Every model, asked to compare a family dessert stop against a named local competitor, brought up Loo.koo.mas. No model, asked the open-ended awareness question, did. That's a specific, actionable finding: the brand wins when the category narrows to Greek donuts or when it's set against a named rival, and it disappears the moment somebody asks the broad question. What that argues for is content and third-party coverage tying the shop to "dessert in Reykjavík" generally, beyond loukoumades.
The pooled 26.7% would never have told you that. Neither would a single headline score.
What a check like this costs
The 15 answers in this scan cost $2.14 in model and search spend, about 14¢ per grounded answer on a premium three-model panel. That's the real invoice rather than an estimate, and it's why credits are denominated one-to-one with answers: one credit, one prompt, one model, one run. What a scan actually costs breaks the arithmetic down per model.
Do it yourself in the next ten minutes
- Write five non-branded questions a customer would type before they know your name. Include one comparison against a rival you'd lose to.
- Run the free checker (two questions, three grounded models, no login) and read the raw answers, not the score.
- For each answer, ask: is my brand named? Is my domain cited? Would a stranger reading this answer end up on my site?
- Note the confidence interval next to the presence rate. If it spans more than about 20 points, you have a direction and nothing more.
- Only then widen to 25+ prompts across all five intent stages, which is where the interval starts to tighten enough to report.
FAQ
How many questions do I need to check whether AI mentions my brand? Five questions across three models gives you 15 observations and an interval roughly ±20 points wide: enough to spot a total absence, not enough to report a rate. For a number you can put in a deck, aim for 25 or more prompts across intent stages, and pool at the topic level rather than hammering a single question.
Should I include my own brand name in the questions I test? Only in a small, clearly-labelled branded set for reputation monitoring. If a question names your brand, the model will almost always mention it, so counting those answers in your presence rate inflates it. Non-branded research and comparison questions are where visibility actually gets won or lost.
Does this measure what people see in the ChatGPT app? No. A scan probes grounded model APIs with live web search. The ChatGPT web app, Google AI Overviews, Google AI Mode and Microsoft Copilot are consumer surfaces with no API, and no model-API measurement substitutes for them.
Why did my brand appear in the answer text but not count as a mention? A mention is only confirmed when an anchor is present: your domain cited in the answer, a category term in the same passage, or a known competitor co-occurring. Answers that name you without one get flagged low-confidence and kept out of the hard metrics, because that's the same pattern a same-named entity produces.