GeoCited

How to Monitor Brand Mentions in AI Answers

Pick 20 to 40 questions your buyers actually ask, run the same list on the same day every week across ChatGPT, Gemini, Perplexity and Google AI Mode, and write down whether you were named, where you ranked and which sites got cited. Compare four weeks, not two. A one week drop of a few points is almost always noise.

Lourdes Paul Agilan, founder of GeoCited

Founder, GeoCited · Published · 6 min

How to monitor brand mentions in AI answers: a frozen prompt set, a fixed weekly run, what to log, and a rule for telling a real change from noise.

On this page

That is the whole routine. The rest of this page is why each part matters, and how to tell a real change from normal wobble. This is the watching job, not the fixing job.

Build a prompt set you promise not to touch

The most common mistake in ai visibility monitoring is rewriting your questions between runs. Do that and you are measuring your own editing, not the market.

Here is how much wording matters. Peec AI, which sells a tracking tool, so treat this as vendor research, ran 37,804 AI responses across 1,754 prompts on five engines in June 2026. Answers held steady while prompts stayed close in meaning. Once wording drifted far enough apart, visibility fell by 2.40 percentage points, about half. Asking for a list or a ranking also got brands named about 20% more often than an open ended question.

Academic work agrees. Kazem Faghih and five co-authors tested 13 models across four benchmarks in May 2026 and found models flipped between right and wrong answers on reworded versions of the same question, mismatch rates above 23%. One prompt tells you very little on its own.

So write your list once. Mix the types: a plain "best X in Y", a budget version, a niche version, a head to head against your main rival, and one question a buyer asks late in the process. Niko Alho, an independent SEO operator in Finland, puts the floor at six to ten prompts. Twenty to forty is better if you sell more than one thing. Then freeze the list. If you must add a prompt, add a new row rather than editing an old one.

Run it on the same day, at the same time

Paul Tschisgale and Peter Wulff, at two German research institutes, sent the same question to GPT-4o every three hours for nearly 88 days, 6,930 queries in total. About 20% of the variation was periodic, with clear cycles at 7.3 days and 5.5 days. The gap between the best and worst point of the cycle was about 14% of the scale, and there was no long term drift.

Be careful with that: it was one physics question on one model, not a brand question. But the lesson is cheap. If hour and weekday can move a score on a fixed question, run your set at a fixed hour on a fixed weekday so time is not one of the variables.

Weekly suits most businesses. Monthly is fine if you publish slowly. As Alho puts it, a snapshot 90 days old is history, not a measurement.

What to log

One row per answer, not per week. Columns:

  • Date and time
  • Engine
  • The prompt, copied exactly
  • Were you named, yes or no
  • Where in the answer you appeared, first, middle or last
  • Which competitors were named
  • Which domains were cited
  • One line on how you were described

Keep the raw answer text in a second sheet. In six weeks you will want to read what actually changed, and a score cannot tell you that.

The citation column is the one that pays. It tells you which pages the engines lean on, and those pages you can go and influence. The yes or no column gives you the score. The citation column gives you the job.

When is a change real

This is the part everyone gets wrong.

Dmitrij Zatuchin, at a university in Tallinn and also at a company selling in this market, published a study in September 2026 on how many times you have to ask before the brand list stops growing. He put 50 questions through six engines 15 times each, 4,500 responses, 1,470 organisations. A single run showed only 62% to 77% of the brands that five runs found. On engines answering without web search, new brands were still appearing at run 15. The median question turned up 38 different organisations across the six engines, and 15 of those appeared on one engine only.

Read that again, because it settles the argument about brand visibility in ai search. If one run finds roughly two thirds of the brands five runs find, then you vanishing from one run on one engine is not news. It is the coin landing tails.

So here is a workable rule. Treat a change as real only if it does at least two of these three things: holds for two weeks running, shows up on more than one engine, and is bigger than the swing you already see in your quiet weeks.

That last one means you need a baseline. Run your set for four weeks while changing nothing on your site. Whatever range you see is your noise floor. Anything inside it is weather.

Single scores mislead. Brand24, a monitoring vendor, pulled 46,350 online mentions between 1 March and 10 May 2026 to find what people complain about here: one prompt on one platform is a lottery ticket, there is no agreed method, and ai brand visibility scores can rise while traffic and revenue fall.

Do you need a tool

Honest answer: probably not at first, and the numbers tools give you are less comparable than they look.

David Nelson ran the same branded prompt through six visibility tools in August 2026. Every one said he was mentioned. The four that put a visibility percentage on it gave 100%, 90%, 75% and 16.7%. His line is worth keeping: a number labelled "visibility" is not automatically measuring the same thing from one product to another.

Prices, read off the vendors' own pages this week. Otterly.AI publishes plainly: $29 a month for 15 prompts, $189 for 100, $489 for 400, enterprise from $1,000, all with daily runs across ChatGPT, Google AI Overviews, Perplexity and Copilot, with Claude, Gemini and AI Mode as paid add ons. Peec AI publishes its plan sizes, 50, 150 and 350 prompts on three models with daily runs, but the prices did not load for me. Profound publishes no price at all; its free trial is 10 prompts run once on ChatGPT only, and the enterprise plan covers up to nine engines. Semrush has folded AI search tracking into its main plans, which start at $139 a month paid monthly.

Where a spreadsheet wins: under about 100 prompts, non English markets, one off audits, and any situation where you want to read the actual answers rather than a score. Alho's do it yourself route runs the queries through an API and costs a few dollars a month for a 36 answer run. That is his figure, not one I could check at source, but the shape is right: the querying is cheap, the dashboard is what you pay for.

Where a tool wins: more than 100 prompts, daily runs, lots of engines, a team that needs a shared dashboard, or history you did not think to collect. We wrote a fuller comparison of buying a tool versus hiring help in tool or agency.

What we will not do is hand you a ranked list of the ten best tools. Those posts get cited and then quietly ignored, because everyone knows who wrote them and why.

Start this week

Write your questions today. Run them tomorrow. Log every answer. Do it again next week at the same hour, and change nothing in between. Four weeks from now you will know your noise floor, and every number after that means something. If you want a second pair of eyes on the first run, our audit covers this.

Sources

  • Peec AI prompt variance study, reported by Search Engine Journal, 15 June 2026. Vendor research.
  • Faghih, Cheng, Saha, Pournemat, Gerami and Feizi, Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy, arXiv 2607.22554, May 2026.
  • Tschisgale and Wulff, Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research, arXiv 2602.15889.
  • Zatuchin, Sampling Completeness in Generative Search: Brand and Cited-Domain Accumulation under Repeated Queries, arXiv 2609.05059, September 2026. Author is affiliated with a company selling in this market.
  • Brand24, Top AI Visibility Tracking Issues 2026 report. Vendor research.
  • David Nelson, Can You Trust AI Search Visibility Scores? I Tested the Same Prompt Across 6 Tools, marketingwithdave.com, 26 August 2026.
  • Niko Alho, How to Measure AI Share of Voice, nikoalho.fi.
  • Otterly.AI pricing page, read 19 September 2026.
  • Peec AI pricing page, read 19 September 2026.
  • Profound pricing page, read 19 September 2026.
  • Semrush pricing page, read 19 September 2026.

Frequently asked

How many prompts is enough to monitor brand mentions?

Six to ten if you sell one thing. Twenty to forty if you sell several, or serve several buyer types. For most small teams more prompts beats more reruns of the same prompt, because you learn about more of your market.

Should I check daily?

No, unless something is on fire. Daily numbers move enough on their own that you will react to nothing. Weekly gives you a comparison you can trust, monthly is acceptable if you publish slowly.

Why do two tools give me different visibility scores?

Because they ask different questions, different numbers of times, on different engines, and add them up differently. One test in August 2026 got figures from 16.7% to 100% for the same brand and prompt. Pick one method and stay with it. Never mix two sources into one trend line.

My mentions dropped 10% this week. What do I do?

Nothing yet. Check whether it held the next week and whether it shows on more than one engine. If it did both, look at the cited domains in your log rather than your own pages, because usually the source the engine leans on changed, not you.

Written by

Lourdes Paul Agilan, founder of GeoCited

Founder, GeoCited

Lourdes Paul Agilan runs GeoCited, a generative engine optimization agency. He works with B2B companies on how AI assistants describe and recommend them. He writes about that work here, and shares the experiments behind it. The tests that worked, and the tests that did not, written up the same way.

LinkedIn

Related reading

Find out what AI says about you.

We run your buyers' real questions against the major answer engines and write down every citation.