GeoCited

How We Pick the Prompts for an AI Visibility Audit (and How We Run Them)

An AI visibility audit is only as good as its prompts. We only test questions where a buyer asks the AI who to hire or what to buy, because that is the only time it names a business. We build those prompts from real Google search terms, test each one before it counts, run the whole set under fixed conditions, and record five separate things per answer.

Lourdes Paul Agilan, founder of GeoCited

Founder, GeoCited · Published · 13 min

How we choose, test and run the prompts in an AI visibility audit, with our own numbers from a 100-prompt ChatGPT run on GeoCited.

On this page

Most AI visibility reports you will see start with a list of prompts someone made up that morning. Then they show you a score.

The score is the easy part. The prompt list is where the audit is won or lost. Pick the wrong questions and you will measure something real, very carefully, that no buyer ever asks.

So this post skips the score. It walks through exactly how we choose the prompts, how we check them, and how we run them. We will use our own numbers from a 100-prompt run on ChatGPT, done on ourselves first.

Why we only test "who should I hire" questions

Think about what you actually type into ChatGPT.

If you ask "what is GEO," you get an explanation. No business gets named. There is nothing to win.

If you ask "best GEO agency for MSPs," you get a shortlist. Names, a first pick, and links to the pages that made the case.

That second kind of question is the only place where an AI recommends a business. So it is the only kind we put in an audit.

Our 24 September run backs this up. We asked ChatGPT 100 buying questions. All 100 answers named at least one business. Between them, they named 280 different agencies. That is a lot of shortlists, and a lot of places you can either be or not be.

Here is what counts as a buying question in our audits:

  • Best or top X for Y: "best msp seo agency", "top agencies for it services companies"
  • Who can do this for me: "which agency can get my msp recommended by chatgpt"
  • X in a place: "best b2b seo agency in london"
  • X for a niche: "best marketing agency for netsuite partners"

And here is what stays out: definitions, how-to questions, tool comparisons with no buying intent, and news. They might be good blog topics. They are bad audit prompts.

Step 1: Start from what people type into Google

This is the part most people skip, and it is the part we care about most.

We do not start by guessing what buyers ask ChatGPT. We start from what they already type into Google. Our source is Google Keyword Planner, checked against the live Google results page.

Why Google, when we are testing ChatGPT? Because of how ChatGPT answers.

Before it writes an answer, ChatGPT runs its own web searches. These are called fan-out queries. We explain the full pipeline in how ChatGPT searches now, but here is the part that matters for picking prompts.

In our 100 runs, we could see 160 fan-out searches. ChatGPT reworded almost every one:

  • It added the word "agency" to 124 of the 160
  • It added "B2B" to 77
  • It added "2026" to 28

But the main term always survived. If the buyer said "MSP SEO," the search said "MSP SEO." If they said "GEO agency," the search said "GEO agency."

So when we start from real Google search terms, we know the AI will be searching the same topics real buyers search. We are not testing a question nobody asks.

How we build the keyword list

1. Make seed terms. We cross three lists with each other.

  • Buyer types: MSP, IT services, software development company, tech company, ERP / NetSuite / Salesforce partner, B2B
  • Services: SEO, GEO, AEO, AI SEO, LLM SEO, marketing, lead generation, content marketing
  • Words people add: agency, company, consultant, services, expert, firm

"MSP" plus "GEO" plus "agency" gives you "MSP GEO agency." Do that for every mix and you get a long seed list fast.

2. Run them through Keyword Planner. Paste the seeds into "Discover new keywords." Set the location to the country your buyers are in (we do the US first, then UK, Canada and Australia) and the language to English. Export everything.

3. Keep buying terms only. A keyword stays if it names a type of provider (agency, company, services, consultant, firm, expert) or has buying words in it (best, top, hire, for [buyer type], in [place]). Anything about definitions, how-tos, jobs, courses, tools, templates, or with "free" in it goes.

4. Check the Google results page. For each keyword you kept, look at page one. If it is full of agency service pages, directories and "best X" lists, it is a buying search. If it is Wikipedia, guides or YouTube, people want to learn, not hire. Drop it.

5. Keep the small ones. This one surprises people. New categories like GEO, AEO and LLM SEO often show 0 to 10 searches a month in Keyword Planner. That does not mean nobody wants them. It means the category is young. If page one looks like a buying page, we keep the keyword and tag it "emerging."

6. Tag every keyword. Buyer type, service (SEO or GEO), place if there is one, monthly searches and cost per click. A high cost per click is a good sign. Advertisers only pay a lot for clicks that turn into customers.

The result is one sheet, one row per keyword. We aim for 200 or more buying keywords before we write a single prompt.

Step 2: Turn each keyword into a question a buyer would ask

Now we turn keywords into prompts. Every prompt asks the AI to recommend a business, and every prompt is built on a keyword from Step 1.

We use six simple templates. Here each one is next to a real prompt from our 24 September set:

Template Real example
best [keyword] best msp seo agency in the us
best [keyword] for [buyer type] best geo agency for managed service providers
top [keyword] that [result] top msp seo companies that get managed service providers ranking on google
which [provider] can [result] which agency can get my msp recommended by chatgpt
best [keyword] in [place] best b2b seo agency in london
best [keyword] for [buyer type] that [condition] best agency for msps that serve healthcare and law firms to rank in ai search

A few rules for writing them:

  • Write it the way a buyer types. Lower case, short, a bit messy. Nobody types a perfect sentence into ChatGPT.
  • No brand names. Not yours, not a competitor's.
  • Never put the client's name in the prompt. A buyer who already knows your name does not need AI to find you. The whole point is to test whether a stranger would hear about you.
  • No leading questions. "Is GeoCited the best GEO agency?" tells you nothing useful.

The test every prompt has to pass

Before a prompt joins the set, we run it once. It only counts if all three of these are true:

  1. The answer names at least one business. If it doesn't, it is not a buying question, no matter how it looks.
  2. The businesses are the kind the prompt asked for. This is the one that catches people out.
  3. The answer reads like a recommendation. A shortlist, a first pick, or "I'd go with." Not a definition.

On our 24 September run, 3 of 100 prompts failed the second test.

Prompt 83 asked for an agency to get a NetSuite partner recommended. ChatGPT listed NetSuite partners, not agencies. It read "NetSuite partner" as the thing we wanted to buy.

The fix is usually small: add the provider word. "Marketing agency that works with Sage partners" works where "marketing agency for Sage partners" gets misread.

A prompt that fails gets rewritten or dropped. It never goes in the set just because we liked it.

Balance the set

If 80 of your prompts are about one buyer type, you learn a lot about that one and nothing about the rest. So we spread prompts across every buyer type and service.

Our 24 September set of 100 looked like this:

  • By service: 51 phrased as GEO, 49 phrased as SEO
  • By buyer type: MSPs 30, IT services 15, tech companies 15, software development 15, platform partners 10, places 10, general B2B 5

The GEO and SEO split turned out to matter a lot. More on that below.

The result of Step 2 is a prompt sheet. Each row has a prompt ID, the prompt, the keyword it came from, buyer type, service, place, the style of wording, and pass or fail on the test.

Step 3: Run every prompt the same way

A good prompt run badly still gives you bad data. So we fix the conditions, write them down, and keep them the same every time.

The conditions we hold fixed

A fresh session every time. In ChatGPT that means a temporary chat, with no memory and no chat history. Your own past chats change what ChatGPT says to you.

Web search on. Location off, or the same on every run. This one bit us. On our first run, ChatGPT used our Chennai location in 30 of the 100 answers. Those answers pulled in local Indian agencies. GeoCited was named in none of them. If we had not caught that, we would have blamed our pages for a problem caused by where we were sitting.

Same model, same day, same machine. We note the model (gpt-5-6 on 24 September), the date and the country of the connection.

Three runs per prompt. ChatGPT's answers change from one run to the next. One screenshot is a sample of one. So we report each prompt as a ratio, like "named in 2 of 3," never as yes or no. We should be honest here: our first run on ourselves used one run per prompt. We now do three, and we are telling you because the numbers in this post come from that single run.

One prompt at a time. Each answer finishes fully before we pull anything out of it.

What we record for every answer

This is where most audits blur things together. We keep five things apart, plus a screenshot:

What we record What it means Where it comes from
Mention Any business named in the answer, and where it was named (1 = first) The answer text
Recommendation A business the answer picks ("top pick", "I'd start with") The answer text
Citation A link shown inside the answer The links in the answer
Source Every page ChatGPT read while searching, linked or not The search behind the answer
Fan-out The exact searches ChatGPT ran The search behind the answer
Screenshot The full answer, the way a buyer sees it The browser

Why keep them apart? Because they tell different stories.

On prompt 7, ChatGPT named GeoCited first, with no link to our site. That is a mention with no citation.

On prompts 28, 52 and 96, ChatGPT read a page on geocited.io while it searched, and then did not name us at all. That is a source with no mention. Our page got in the room and did not get picked.

Those two problems need different fixes. If you merge them into one "visibility score," you can't tell which one you have.

Here is how our own numbers came out across the 100 prompts:

Measure Result
geocited.io read as a source 8 prompts
GeoCited named 5 prompts (first in 3)
geocited.io linked 4 prompts
Named on SEO-worded prompts 0 of 49
Links in all answers 574 (5.7 per answer)
Links that went to "best X" lists 315 (55%)
Answers with a clear top pick 31

Every one of our five wins came from a GEO-worded prompt. On the 49 SEO-worded prompts we got nothing. That was fair. Our pages are written about AI search, not general SEO, so ChatGPT had nothing on our site to match those prompts to. Until we publish SEO pages of our own, those 49 prompts are our control group: if our numbers on them move, something outside our work changed the results.

How our internal tool runs all of this

We did our first 100-prompt run by hand in the browser. Type a prompt, wait, pull the full conversation record, take a screenshot, move on. It worked, and it gave us the numbers above. But it is slow, and it does not scale to three runs per prompt for every client.

So we built an internal auditor that does Step 3 for us.

You give it the business name, the website and the prompt sheet from Step 2. It sends every prompt to ChatGPT's search mode through a data service, not through someone clicking in a browser. Then, for every answer, it stores:

  • The full answer text
  • Every link in the answer, in the order it appeared
  • Every page ChatGPT read while searching
  • The fan-out searches, when they are exposed
  • Every business named, and whether it was named as a pick

From that, it builds the report on its own:

  • Share of voice. Who got named, how often, and how often first.
  • A mention ledger. Every business, on every prompt, with its position.
  • Page-level citations. Not just which websites got linked, but which exact pages.
  • Splits by buyer type and service, so you can see MSP prompts apart from IT services, and GEO apart from SEO.
  • Audit-to-audit comparison, so a re-run at day 30 sits next to the baseline.

It also exports everything as spreadsheets, so nothing is locked inside it.

The rule we built it on is simple: it never makes a number up. Every mention, link and search in the report traces back to the raw answer it came from. If a figure is worked out rather than read directly, like a link's position, the report says so. And when ChatGPT does not expose its fan-out searches for an answer, the tool marks that answer as "not shown" instead of pretending it ran zero searches.

That is the whole point of running it this way. You should be able to click any number and see the answer behind it.

What we do with the data

Running the prompts gives you data. The next step turns it into a to-do list. We will cover that in full in a follow-up post, but here is the shape of it, so you can see why the prompt work matters.

Every fix must point to a prompt we lost. For each lost prompt, we note which business won it, which page carried the win, and what we are missing. If we can't point to a lost prompt, the fix doesn't go on the list. No "best practice" filler.

We read how each answer was built. If more than half its links go to "best X" lists, it was built from lists, and the fix is getting onto those lists. If not, it was built from company pages, and the fix is on your own site. In our run, 55% of all links went to lists. We were on only 1 of the 12 most-linked lists and directories, and that one listing gave us our first-place mention on prompt 7.

We re-run the exact same prompts. Same set, same conditions, at day 30, 60 and 90. A change only counts as working when the same prompts, run the same way, name you more often. Everything else is a guess.

The rules we follow on every audit

These came from mistakes and surprises on our own first run.

  1. Buying questions only. If a prompt names no business, or names the wrong kind, it doesn't count.
  2. Never trust one screenshot. Report ratios from three runs.
  3. Location off, or held the same. It changed the sources in 30% of our answers.
  4. Mentioned, read and linked are three different numbers. Never merge them.
  5. Record the fan-out word for word. The words ChatGPT adds explain why some pages get found and others don't.
  6. Open the linked pages yourself. Don't assume you are or aren't on a list. We checked 12 by hand.
  7. Every fix names its prompts.
  8. Keep a control group. Prompts you have no page for should not move. If they do, something outside your work changed.
  9. Keep the client's name out of the prompt.
  10. Use the buyer's exact words on the page. The fan-out keeps the main term and adds its own. Every winning page we saw carried that exact term in its title and main heading. Whether that is the reason it won is what our next re-run will test.

Want us to run this on your business?

If you want to see what ChatGPT says when your buyers ask who to hire, we run this method as our AI visibility audit. You get the prompt sheet, the raw answers, and every number traced back to the answer it came from.

Or take the steps above and run it yourself. The method is the same either way. The tool just saves you the hours in the browser.

Sources

  • GeoCited, 100-prompt ChatGPT audit, run 24 September 2026 (gpt-5-6, temporary chat, web search on, one run per prompt, from Chennai, India). Raw answers, sources, fan-out and screenshots held internally.
  • GeoCited, ChatGPT citation gap analysis, 25 September 2026, based on the same run.
  • Google Keyword Planner, United States, used for the keyword list behind the prompt set.

Frequently asked

What is an AI visibility audit?

An AI visibility audit checks whether AI assistants like ChatGPT name your business when buyers ask who to hire or what to buy. It runs a fixed set of buying questions, records who gets named, linked and read as a source, and shows which pages decided each answer.

How do you choose prompts for an AI visibility audit?

We start from real Google search terms in Keyword Planner, keep only buying terms, turn each one into a question that asks for a recommendation, and test each question once. A prompt only counts if the answer names businesses of the type it asked for.

Why not just make up the prompts?

Made-up prompts often test questions nobody asks. ChatGPT runs its own web searches before answering, and in our 100 runs those searches kept the buyer's main term every time. Starting from real search terms means you test what buyers actually look for.

How many prompts do you need?

We aim for about 100 in the first run, spread across every buyer type and service. For tracking change over time we freeze around 60, including about 10 control prompts that should not move.

How many times should you run each prompt?

Three times, under the same conditions. ChatGPT's answers change from run to run, so a single answer can't tell you much. Report each prompt as a ratio, like named in 2 of 3.

What is the difference between a mention and a citation?

A mention is your business name in the answer. A citation is a link to a page in the answer. You can be named with no link, and a page of yours can be read as a source without you being named. They need different fixes.

Why turn location off?

Because ChatGPT will use it. In our run it used our location in 30 of 100 answers and pulled in local agencies. If your buyers are in another country, a location-based answer is not what they see.

Written by

Lourdes Paul Agilan, founder of GeoCited

Founder, GeoCited

Lourdes Paul Agilan runs GeoCited, a generative engine optimization agency. He works with B2B companies on how AI assistants describe and recommend them. He writes about that work here, and shares the experiments behind it. The tests that worked, and the tests that did not, written up the same way.

LinkedIn

Related reading

Find out what AI says about you.

We run your buyers' real questions against the major answer engines and write down every citation.