Skip to content

AI Search

How to measure whether your brand appears in AI answers

Build a fixed set of prompts covering how people describe your category, your brand, and the topics you claim expertise in. Run each in a clean session with memory and personalisation off, several times, across the assistants your buyers use. Score every result as cited, mentioned, or absent, and record which competitors appeared instead. Repeat on a fixed schedule.

By Viken Patel

Most companies have no idea how they appear in AI-generated answers. Not a rough idea. No idea at all, because nobody has looked in a structured way. The checking that does happen is usually one person typing their company name into ChatGPT once and forming an impression from a single response.

That is not measurement. Here is a method that is.

Build a fixed prompt set

Three tiers, each answering a different question.

Entity prompts establish how the models classify you. "Who is [company]?" "What does [company] do?" "Is [company] a [category A] or [category B] business?" These reveal whether the model's understanding of you matches what you think you are known for. In my experience this is where the most uncomfortable findings come from.

Commercial prompts are the ones with money attached. "Who can help my company with [problem]?" "Best [category] providers for [segment]." "I need [service]. Who should I consider?" These are the prompts your buyers actually run, and being absent from them is a pipeline problem, not a marketing one.

Topic prompts test whether your content is cited on the subjects you claim. "How do I [task you write about]?" "What is the difference between [two concepts you cover]?"

Fifteen to twenty-five prompts total is enough. More is not better, because you have to run them all again next quarter and an oversized set gets abandoned.

Write them once, then freeze them. If a prompt turns out to be poorly worded, retire it and add a replacement rather than editing it in place. An edited prompt makes every future comparison invalid.

Run them under controls

Get any of these wrong and you have measured nothing.

Clean sessions. Logged out, or a temporary chat. Memory off. Custom instructions off. One prompt per fresh session, with no follow-ups in an existing thread.

This is the control people most often skip, and it is the one that most reliably produces false results. If you ask an assistant about your company from your own account with memory enabled, it may be drawing on what it knows about you as a user (your previous conversations, your stated occupation) rather than what it knows about your company as an entity in the world. People come away reassured by a result that reflects nothing but their own chat history.

Repeat runs. These systems are non-deterministic. The same prompt produces different answers. One run is an anecdote. Run each prompt two or three times and record every result, because the frequency is the finding: "cited in two of three runs" is real information that a single check cannot give you.

Consistent conditions. Same location, same language, every round. Record whether browsing was enabled and keep that consistent per surface.

Record the model version. Model updates will explain more variation in your results than your website changes will, and you need to be able to tell the difference.

Separate retrieval from training

This distinction determines what your results mean, and leaving it out is the most common analytical error I see.

With browsing disabled, an assistant answers from what it learned during training. That is a snapshot of the web from some point in the past. Nothing you do to your site changes it until a future training run, which for most companies may never meaningfully happen.

With browsing enabled, the assistant searches, retrieves, and reads live. That reflects your site as it is now, and it is the surface your work actually influences.

Test both. Report them separately. If you combine them into one number, you will either conclude that your work had no effect when it did, or claim an effect you did not produce.

Score it consistently

A simple three-point scale is sufficient:

Cited (2). Your domain appears as a linked source. Mentioned (1). You are named in the answer without a link. Absent (0). Neither.

Average across runs, then across prompts, and you have a comparable score. The absolute number means little. The direction of travel between rounds is the whole point.

Record more than the score

Three fields matter as much as the number:

Which URLs were cited. Tells you which content is doing the work, which is usually not what you expected.

How you were described. This is the single most valuable field in the whole exercise. The exact words the model uses to characterise your company reveal how the entity is currently understood. When that sentence changes over time, you have evidence of repositioning that no other metric captures.

Which competitors appeared instead. On commercial prompts, absence is only half the finding. Knowing who occupied the space is the other half.

Set a cadence and expect slow movement

Baseline, then thirty, ninety, and one hundred and eighty days.

Be realistic about what moves and when. Retrieval-based surfaces can shift within weeks of a meaningful site change. Entity descriptions shift slowly, over months. Training-derived answers may not shift at all in any timeframe you care about.

A consultant promising to change how ChatGPT describes your company within a month is either not distinguishing these surfaces or is not being straight with you.

The part that makes it credible

Decide what success looks like before you see any results, and write it down.

This sounds like an academic nicety. It is the difference between a measurement and a story. If you have not defined the target in advance, you will find yourself, six months later, quietly selecting whichever number happened to improve and building the narrative around that. Everyone does this. It is not dishonesty, it is how people reason when the criteria are ambiguous.

Write down the hypothesis, the number that would count as success, the number that would count as failure, and whether you will report the result either way. Then run it.

That last commitment is what makes the eventual result worth anything to anyone else.

This article is part of the SEO in the AI Era: The Complete Guide guide.

FAQ

Questions this raises

Why do I get a different answer every time I ask the same question?
These models are non-deterministic, so identical inputs produce varying outputs. This is why a single check tells you very little and why each prompt needs to be run several times, with the frequency of citation recorded rather than a single yes or no.
Can I use the API instead of doing this manually?
You can, but the results will not match what your buyers see. The API and the consumer product are different systems with different retrieval behaviour and different system prompts. For measuring buyer-facing visibility, test the product your buyers actually use.
How often should I re-run this?
Quarterly is enough for most companies. Run it more often only around a significant site change, and always with identical prompts, or the comparison is meaningless.