2026-06-22 · 10 min read · Practitioner how-to

How to Audit a Competitor's AI Search Visibility

Somewhere in your competitor's marketing team, someone is probably asking ChatGPT for tool recommendations in your category, and their brand is showing up in the answer. Yours might not be. You will not see this in Google Search Console, in your rank tracker, or in any paid ad dashboard. The only way to know is to ask the models the same questions your buyers are asking, and read what comes back.

This guide walks through a manual method for auditing a competitor's AI search visibility: which prompts to run, across which models, how to log what you find, and how to turn a spreadsheet of raw answers into a prioritized list of content and PR moves. It takes a few hours the first time and less on every repeat. At the end, we will also cover why most teams stop doing this by hand after the second or third round, and what daily monitoring looks like instead.

What does an AI visibility audit actually measure?

An AI visibility audit answers one question: when a prospective buyer asks an AI assistant about your category, does a given brand get mentioned, and how favorably? That is different from traditional competitive analysis, which looks at keywords, backlinks, and rankings. Here you are looking at whether a brand is cited, recommended, or compared inside a generated answer.

Three things matter for each prompt you run:

You are effectively reverse-engineering a black box. ChatGPT, Perplexity, Gemini, and Claude do not publish a ranking algorithm you can inspect. What you can do is sample it repeatedly, the same way you would test any system you cannot see inside: same prompt, same models, tracked over time, with enough repetitions to separate a pattern from a fluke.

Step 1: Build a prompt list that mirrors real buyer questions

The single biggest mistake in a manual audit is testing prompts that flatter your hypothesis instead of ones a real buyer would type. Skip "what is the best {category} tool" as your only input. Buyers phrase things with more context, more constraints, and more comparison intent than that.

Build your prompt list from four categories:

Aim for 15 to 25 prompts for a first pass. Fewer than that and you will not see a pattern; more than that and the manual logging becomes the bottleneck before the insight does. Keep the prompt wording identical across every model run, and write it down. Small rewording changes what gets retrieved, so consistency matters more than cleverness here.

Step 2: Run the same prompts across every model your buyers actually use

Do not audit a single model and assume the result generalizes. ChatGPT, Perplexity, Gemini, Claude, and Copilot each pull from different retrieval systems, weight sources differently, and browse the web on different schedules. A brand that appears reliably in Perplexity answers can be entirely absent from Gemini for the same query, and the gap itself is useful information.

A few things worth knowing about how these systems actually work before you start logging results:

ChatGPT's search decides, per query, whether the moment calls for a live web lookup or not, then surfaces inline citations you can click through to the source page, according to OpenAI's own documentation on ChatGPT search. That means the same prompt asked twice, minutes apart, can trigger different retrieval behavior depending on how the model reads your intent. Run each prompt more than once before concluding a competitor is or is not present.

Practically, for each of your 15 to 25 prompts, run it fresh (new chat, logged-out state where possible, no prior conversation context) in:

That is roughly 75 to 125 individual queries for a full first pass. Budget two to three hours if you are doing it properly, longer if you are logging framing and sentiment in detail rather than a simple yes or no.

Step 3: Log every answer the same way

A spreadsheet beats a document here, because you need to sort, filter, and compare across models later. One row per prompt per model. Columns worth including:

The sources column matters more than people expect. If a competitor keeps getting cited off the back of the same three articles (a review roundup, a Reddit thread, their own comparison page), that tells you exactly where to focus outreach or content. This lines up with what Search Engine Land's guide to competitive audits for AI search recommends: pull the specific page an AI answer is drawing from and compare it side by side against your own equivalent content, rather than treating the citation as a black box.

How do you turn raw answers into a score?

Once the spreadsheet is full, convert it into something a stakeholder can read in thirty seconds. A simple scoring model works better than an elaborate one, because you will want to repeat this monthly and compare over time.

A workable rubric per prompt, per brand:

Average the score across all prompts and all models for each brand, and you get a rough "share of voice in AI answers" number you can track release over release. Break it out by model too. A brand that scores well on Perplexity but zero on Gemini has a specific, fixable gap rather than a vague "we need more AI visibility" problem.

If you want the underlying mechanics of why some brands score consistently higher, the seven structural factors we cover in our breakdown of what actually drives AI visibility are the same factors worth checking against your competitor's content, structured data, and citation footprint while you have their answers open in front of you.

A worked example: scoring one prompt across four models

It helps to see the rubric applied before running your own audit. Take a single prompt: "what tools should I use to track brand mentions in AI chatbots." Here is roughly what a scoring pass might look like for one hypothetical competitor, brand names aside:

Average score for that one prompt: 1.5 out of 3. Already this tells you something specific: the competitor's homepage and pricing page are doing real work in Perplexity's citation pipeline, they have essentially no presence in Claude, and Gemini treats them as a follow-up answer rather than a default recommendation. Run the same five-minute exercise across your full prompt list and the pattern across models becomes the story, not any single answer.

What to do with the findings

An audit that ends in a spreadsheet is a wasted afternoon. The findings should route into three places:

Re-run the same prompt list on a fixed cadence, monthly at minimum, so the score is a trend rather than a snapshot. Models update retrieval behavior and underlying indexes constantly, and a result from March tells you very little about June.

Where the manual method breaks down

The approach above works. It also does not scale past a handful of prompts and a monthly cadence, for reasons that show up the second time you try it rather than the first.

Model behavior shifts week to week, not just month to month, so a monthly snapshot misses the week your competitor's new comparison page starts getting cited and the two weeks after that when you could have responded. Logging framing and sentiment by hand is subjective and inconsistent between whoever runs the audit each cycle. And the moment you want to track more than two or three competitors across five models and twenty prompts, you are looking at 300-plus manual queries every cycle, which nobody sustains past quarter two.

This is the exact gap daily, automated tracking is built for. If you want to see what the output of a structured check looks like before committing to a manual process, our guide to checking your AI brand visibility walks through the same core prompts-and-scoring logic in a version you can run in minutes.

What Bold GEO automates from this exact process

Bold GEO runs the method above on a daily refresh instead of a manual, occasional one. We track how your brand and your named competitors are cited across ChatGPT, Perplexity, Gemini, Claude, and Copilot, using a consistent prompt set tied to your category, then log presence, position, and framing automatically so you get the scoring rubric above without the spreadsheet.

Because the checks run daily, you see a shift the week it happens, not the quarter after. A competitor's new comparison page picking up citations, a model updating which sources it trusts, a sudden framing change after a product launch: all of it shows up as a trend line instead of something you would have needed to catch manually, mid-cycle, by re-running the prompts yourself.

You can start on Bold GEO's daily AI visibility tracker with a 7-day free trial, no credit card required, or run individual scans pay-as-you-go at $1 per scan if you want to test the waters on a specific competitor set before committing to a monthly plan.

A simple way to start this week

You do not need every model, every competitor, and a formal rubric on day one. Pick your three most important buyer-facing prompts, the ones closest to a purchase decision, and run them across ChatGPT and Perplexity today. Log presence and position for your brand and your top competitor. That alone will tell you whether you have a real gap worth investigating further.

From there, expand the prompt list, add the remaining models, and set a recurring calendar reminder to re-run it monthly, or hand the recurring part to a tool built to do it daily. Either way, the visibility gap is only invisible until you decide to look.

Track your brand in AI answers. Start free.

Bold GEO monitors how your brand is cited across ChatGPT, Perplexity, Gemini, Claude, and Copilot on a daily refresh. 7-day free trial, no credit card.

Start free trial →