How to Improve Your AI Visibility (and How to Measure It First)
To improve your AI visibility, measure it first. Pick 20 to 40 questions a real buyer would ask, run each one several times on ChatGPT, Gemini and Perplexity, and write down whether you get named. Repeat weekly. Then fix what the log shows: more mentions of your brand on other websites, clearer pages, and content that answers those exact questions.
How to improve your AI visibility: build a repeatable baseline across ChatGPT, Gemini and Perplexity first, then fix what the log shows.
On this page
That order matters. Here is why.
AI visibility is not one number
In August 2026 the IAB published a framework called Measuring Visibility in the AI Era. One line in the announcement is worth sitting with: more than 20 companies now sell AI visibility measurement tools, each using a different method, so the same brand can get different answers from different vendors on the same day.
The IAB splits visibility into four parts. Presence is whether you get mentioned at all. Prominence is where in the answer you land. Portrayal is whether the AI describes you correctly. Persuasion is whether any of it makes someone act. Most tools sell you Presence and call it visibility.
The framework also draws a useful line. Some measurement is directional, good enough to show a trend but not to move budget on. Some is decision grade, meaning it meets a bar for sample size and repeatability. Caroline Giegerich, the IAB's VP of AI, put it plainly: people are discovering brands inside AI tools, and measurement has not kept up.
The reason a single check tells you nothing
If you ask ChatGPT a question today and it names you, that is not a result. It is one draw from a dice roll.
The cleanest test of this came from the University of St. Gallen. Julius Schulte, Malte Bleeker and Philipp Kaufmann tracked four industries across ChatGPT, Gemini, Google AI Mode and Perplexity for about six weeks in early 2026. Day to day, the sources the AI pulled from overlapped only 34 to 42 percent. Brand mentions were steadier at 45 to 59 percent. Then they ran the same prompts again minutes apart and saw the same instability, so the wobble comes from the model itself, not from anything changing in the world.
A second paper, from Ronald Sielinski in July 2026, asked how much data you need before a ranking of brands stops moving. Across 30 platform and topic combinations on Gemini, SearchGPT and Perplexity, he found it takes between 33 and 94 answers before the order settles. Three of the 30 tests never settled at all, even after 125 questions, and all three were on SearchGPT. Worth noting: his questions came from ChatGPT rather than real searches, and he sells software in this space.
So when a tool hands you a tidy score, ask how many runs sit behind it. If the answer is one run per prompt, the number is noise wearing a suit.
How many runs, then? The experts disagree
The St. Gallen team says run each prompt at least seven times a day for brand tracking, eight for source tracking, and look at a rolling two to four week window. Nick Lafferty, who keeps a public metrics reference updated through August 2026, says the same thing in plainer words: a 50 prompt set run 10 times tells you more than a 500 prompt set run once.
Dmitrij Zatuchin disagrees. In July 2026 he broke 12,933 AI responses covering 20 brands, eight languages and three models into buckets of variation. Repeating the same prompt explained 34.8 percent of the wobble and the language of the question explained 31.6 percent, while the model used explained only 1.7 percent. His conclusion: a repeat past the fifth buys almost nothing, and those queries are better spent on more languages and contexts. He is affiliated with a company in this market, and his brands were all Central and Eastern European, so it is not a like for like test.
Nobody has reconciled the two. My read: five runs per prompt is enough to stop fooling yourself, and anything past ten is money better spent widening your question list.
Do it yourself first
You do not need a subscription to start. Open a spreadsheet.
List 20 to 40 questions a real buyer types, not the ones you wish they typed. Your sales team knows them. Run each one five times in ChatGPT, Gemini and Perplexity, in a fresh chat each time, with memory and personalisation off so you are not scoring yourself against your own history. Log two things: were you named, and were you linked. Repeat next week.
After a month you have a baseline built on a method you understand, and you will read any tool you later buy far better for it. If you would rather someone else run that first pass, that is what a GEO audit is for.
A paper published on 14 September 2026 by Edward Malthouse and colleagues tested brand recommendations across six models and five product categories. Broad category questions left out well established brands surprisingly often, but those same brands often reappeared once the question carried a specific situation. So test both shapes. Ask "best IT support company" and also "IT support for a 40 person law firm that has to stay HIPAA compliant."
The numbers on your own site that nobody can fudge
Prompt logs tell you about the AI. Your own analytics tell you about your business. Wil Reynolds of Seer Interactive argued in February 2026 that visibility scores are a vanity metric until you tie them to something real, and pointed at direct traffic, branded search and where AI visitors land.
Search Console has a generative AI performance report, rolled out worldwide on 31 August 2026. Impressions only, no clicks, and no way to pull AI Overviews apart from AI Mode. Limited, but free and yours.
Your analytics referral report shows visits from chatgpt.com, perplexity.ai and gemini.google.com. Expect the volume to look tiny. Seer's widely quoted case study found AI traffic was 0.07 percent of organic sessions on the one site they studied, though it converted at 15.9 percent against 1.76 percent for Google organic. Read that carefully. One client, 2025 data, and the multiples the industry quotes run from 4x to 23x depending on who is selling. Something is there. The size is unsettled.
Branded search volume is third. If more people hear your name inside AI answers, more of them later type it into Google. Slow signal, but hard to fake.
Now the fixes
Once you have a baseline, the work is less exotic than the tools suggest.
Get mentioned elsewhere. Ahrefs, in vendor research covering 75,000 brands, found branded web mentions had the strongest link to AI Overview visibility of anything they measured, ahead of backlinks. They stress it is correlation, not proof. It points the same way as the Malthouse paper, which found prominence tracked how much a brand was discussed and searched for rather than how big it was.
Answer the exact questions in your log. Not variations. If three prompts keep naming your competitor, write the page that answers those three prompts better than the page the AI is currently quoting. Our guide on how to optimise content for AI search covers the page level work.
Then wait properly. Data from about 900 marketing pages tracked between March and May 2026 put the median time to a first AI citation at about seven days, with a long tail past a month. That comes from a vendor platform, so hold it loosely. Checking on day two proves nothing.
Where to start this week
Build the question list. Run each prompt five times across three platforms. Log named and linked. Open your Search Console AI report and write down today's impressions. Then leave it alone for a week. You now have what almost nobody in your market has: a before.
Sources
- IAB, Measuring Visibility in the AI Era, announcement and framework, 3 August 2026. iab.com
- Julius Schulte, Malte Bleeker and Philipp Kaufmann, University of St. Gallen, Don't Measure Once: Measuring Visibility in AI Search (GEO), arXiv 2604.07585, 10 April 2026.
- Ronald Sielinski, From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement, arXiv 2607.10341, 11 July 2026. Author is a co-founder of IQRush. Headline figures read via Search Engine Journal, AI Visibility Rankings Aren't Stable, 11 July 2026.
- Dmitrij Zatuchin, Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers, arXiv 2607.13304, 14 July 2026. Author affiliated with Rankfor.AI.
- Edward Malthouse, Kun-Yu Lee, Jing Yang, Sanchary Pal and Xueyan Feng, Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations, arXiv 2609.16304, 14 September 2026.
- Nick Lafferty, AI Visibility Metrics: Formulas, Benchmarks and Sample Sizes, published 9 June 2026, updated 20 August 2026. Time to first citation figures drawn from Profound datasets, a vendor source.
- Wil Reynolds, Seer Interactive, AI Visibility Is a Vanity Metric, 3 February 2026.
- Nick Haigler and Garman Chan, Seer Interactive, Case Study: How Traffic from ChatGPT Converts, 3 June 2025. Single client, GA4 data from 1 October 2024 to 30 April 2025.
- Louise Linehan and Xibeijia Guan, Ahrefs, An Analysis of AI Overview Brand Visibility Factors, 26 May 2025. Vendor research, 75,000 brands, Spearman correlation.
- Google Search Console Help, Generative AI performance report, accessed 18 September 2026. Global rollout stated as 31 August 2026.
- Search Engine Roundtable, September 2026 Google Webmaster Report, accessed 18 September 2026.
Frequently asked
How often should I check my AI visibility?
Weekly is enough for most businesses, monthly if your market moves slowly. Daily checking shows you swings that are just the model being random. The St. Gallen research found day to day source overlap as low as 34 percent when nothing had changed.
Do I need to buy an AI visibility tool?
Not to start. A spreadsheet and a few hours give you a real baseline. Buy a tool when the manual work gets too big, and when you do, ask how many times they run each prompt and whether they report a margin of error. The IAB calls anything failing that bar directional rather than decision grade.
Why does my AI visibility score differ between two tools?
Because they measure different things with different prompts and different run counts. The IAB noted in August 2026 that more than 20 vendors each use their own method and can return different answers for the same brand. Pick one tool and track your own trend instead of comparing scores across vendors.
Is AI visibility worth measuring if the traffic is so small?
The traffic is small today, the quality signals look good, and the numbers are contested. Measure it because measuring is cheap, and because branded search and direct traffic will tell you whether it is landing. Do not rebuild your marketing around a percentage nobody can reproduce yet.
Written by

Lourdes Paul Agilan runs GeoCited, a generative engine optimization agency. He works with B2B companies on how AI assistants describe and recommend them. He writes about that work here, and shares the experiments behind it. The tests that worked, and the tests that did not, written up the same way.