Every AI visibility report rests on a list of questions. Change the questions and the numbers change with them, often dramatically.
That makes the prompt set the most important and most neglected part of AI visibility tracking. Teams spend weeks choosing a tool and ten minutes choosing what to ask it.
This guide walks through how we build prompt sets: where the prompts come from, how to structure them, and how to keep the set stable enough that trends actually mean something.
Imagine two marketers tracking the same brand. One asks "what is the best CRM", the other asks "which CRM should a five-person real estate team use". The first will almost never see a niche product mentioned. The second might see it in every answer.
Both numbers are accurate. Only one of them reflects how that brand's real buyers ask. A prompt set is a model of your market, and a bad model produces confident, useless reports.
If you have not yet run a manual baseline, start with our guide on how to check your brand's AI visibility. It will give you a feel for how answers vary before you commit to a structured set.
The best prompts are not invented. They are borrowed from the people you sell to. Before writing a single prompt, gather raw material from at least four places.
Aim for 100 or more raw questions. Most will be duplicates or near-duplicates, and that is useful: the questions that come up again and again are the ones worth tracking.
Group the raw questions by what the person is trying to do. We use five intent groups, because each one tells you something different about your position.
| Intent | Example prompt | What it tells you |
|---|---|---|
| Category discovery | What are the best tools for tracking brand mentions in AI answers? | Whether you are in the consideration set at all |
| Use case | How can an agency report AI visibility to clients? | Whether you are associated with the jobs you are best at |
| Comparison | How does Tool A compare with Tool B for small teams? | How you are framed against named competitors |
| Problem | Why doesn't ChatGPT mention my company? | Whether you are seen as an authority on the underlying problem |
| Branded | What does Bold GEO cost and what does it track? | Whether the facts about you are correct |
A healthy set leans heavily on the first three groups. They are where new buyers are won or lost. Branded prompts are a smaller accuracy check.
People do not type keywords into AI assistants. They write sentences, add context about themselves and often ask two things at once. Your prompts should look the same.
Compare these two versions:
Keyword style: best email marketing software small business
Prompt style: I run a small online store with about 2,000 subscribers. Which email marketing tool is easiest to set up and won't get expensive as I grow?
The second version produces a completely different answer, usually with fewer and more specific recommendations. It is also much closer to what your buyer really types. Keep the context realistic, though. Every added detail narrows the answer, so only include details your typical buyer would actually mention.
If you sell to distinct segments, write the same core question from each point of view. "Best CRM for a solo consultant" and "best CRM for a 50-person sales team" are different battles, and you may be winning one and absent from the other.
There is a tension here. More prompts give you more stable numbers. More prompts also mean more answers to read and more noise to filter.
For most brands we recommend starting with 25 to 60 prompts, spread roughly like this:
Adjust the mix to your market. A brand in a crowded category with strong named competitors should weight comparisons higher. A brand creating a new category should weight problem prompts higher, because buyers do not yet know the category name.
Visibility is not one number. A brand can be a default recommendation in Perplexity and invisible in Claude, because the assistants source answers differently. We explained why in AI Overviews vs AI chat.
Run the full set across ChatGPT, Perplexity, Gemini, Claude and Copilot. Note whether each assistant used live web search for a given answer. OpenAI documents when ChatGPT searches the web (ChatGPT search), and Google explains how its AI features draw on its index (AI features and your website). Answers with search enabled react to changes much faster.
Reading answers is not the same as measuring them. Decide up front what you capture, and capture it the same way every time.
These fields roll up into the metrics that matter, especially share of voice. Our guide to tracking share of voice in AI answers covers the calculation in detail.
The value of tracking comes from comparing today with last month. That only works if the questions stay the same.
Mark 70 to 80 percent of your prompts as core and do not edit them. Even small wording changes can shift results, so treat core prompts like a fixed benchmark.
Use the remaining prompts to test new angles: a new use case, a new competitor, a new persona. When an exploratory prompt proves useful for a full quarter, promote it into the core set and note the date.
AI answers are not deterministic. The same prompt can return different brands on different days. Running prompts on a regular schedule, ideally daily, lets you see trends instead of reacting to a single lucky or unlucky answer.
If you want to start today, use this skeleton and fill it with your own language from Step 1.
Write three to five variants of each and you have a working first set of 25 to 35 prompts. Refine from there as the answers teach you which questions matter.
To show how the pieces fit together, here is how the first draft of a set might look for a cloud accounting product aimed at freelancers and small agencies. The raw material came from sales call notes, support tickets and a freelancer community forum.
Notice the details that came straight from customers: hating spreadsheets, late-paying clients, invoicing in two currencies. Those specifics are what make an assistant's answer match the real buying situation. They are also exactly the details a generic keyword list would have removed.
From here, the team would write two or three variants of each prompt, mark the strongest 25 as core, and start daily runs. After a month they would know which intents they win and which ones competitors own.
A prompt set is a model of how your market asks questions. Build it from real buyer language, group it by intent, write it the way people actually prompt, and keep a stable core so trends mean something.
Get this right and every other part of AI visibility work, from content to outreach, has a reliable scoreboard.
For most brands, 25 to 60 prompts is the practical range. Fewer than 20 makes the numbers jumpy, and more than 100 is hard to review by hand. Grow the set as you learn which intents actually move.
Keep most of them unbranded, because that is where you win or lose new buyers. Add a small branded group to check accuracy: what the assistants say about your pricing, features and positioning when someone asks about you directly.
Change it rarely and deliberately. Every change breaks comparability with earlier results. Review it quarterly, retire prompts that no longer match how buyers search, and note the date of every change next to your trend line.
Bold GEO monitors how your brand is cited across ChatGPT, Perplexity, Gemini, Claude, and Copilot on a daily refresh. 7-day free trial, no credit card.