← All articles

How we run an AI visibility audit: our methodology, step by step

Sepia-toned photo of hands checking results on a laptop
Key takeaways
  • The method in one line: 5 or 30 real buyer questions, run by hand on a logged date across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews, with every answer read and every claim about you checked by a person.
  • What gets scored: mention rate, rank inside the answer, and sentiment, per engine and per question, for you and every competitor the engines named.
  • What a scanner skips: each wrong claim traced to the source that fed it, and a fix card per finding with one owner, an effort estimate, numbered steps, the ready-to-use asset and a "done when" check.
  • What it costs and how long it takes: Snapshot 149 EUR in about 3 business days; Full Audit 690 EUR in 5 to 7 business days with 7 monthly re-measures included; both ex-VAT with a 14-day money-back guarantee.

A buyer who asks ChatGPT which provider to pick gets three or four names and a reason for each, and the brands left out never learn they were in the running. MentionShare's AI visibility audit methodology answers that in a checkable way: a person runs your real buyer questions (5 in a Snapshot, 30 in a Full Audit) through ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews on a logged date. Every answer is read, every claim about you is checked, every wrong claim is traced to the source that fed it, and every finding ends in a fix card. One qualifier applies from the first step: AI answers vary between runs, so a single answer counts as one measurement, and the baseline rests on the pattern across questions, engines and re-runs.

This page walks through the method in the order an audit actually runs, from writing the questions to the seventh monthly re-measure. It then shows what the report contains, how to run a rough version yourself this week, and where the method stops.

What does an AI visibility audit measure?

An AI visibility audit measures two things: whether AI engines name your brand when buyers ask category questions, and whether what the engines say about you is true. Absence and inaccuracy are different failures with different fixes. A brand can be missing from every "best X for Y" answer because no source the engines read records it, and a brand named in every answer can still lose the deal to a wrong price quoted with full confidence. Our audit tests both, engine by engine, because the results differ: a brand can be strong in Claude and absent from Gemini on the same question.

If the category is new to you, our guide to what a GEO audit is covers price tiers and deliverables across providers, and the GEO explainer covers why AI answers work the way they do. This page is the operating manual for our own audit.

Step 1: Which buyer questions do we test, and how do we baseline them?

We start by writing the questions your buyers actually type, because visibility is measured per question and being named on "best CRM" says nothing about "best CRM for a two-person sales team". A Snapshot tests your 5 most important questions and a Full Audit tests 30, written with you after we learn the category. They cover the buyer situations you win, a price question ("how much does [category] software cost"), alternatives questions naming the market leader, and the trust questions that surface late in a decision ("is [brand] legit", "who owns [brand]"). Questions that name your brand test accuracy, and category questions test discovery; the panel needs both, and the category questions carry the share-of-voice measurement because that is where buyers actually meet you.

Each question then runs through all five engines on one logged date, in clean sessions, one question per new chat so an earlier conversation cannot colour the answer. For every answer we record the brands named in order, the reasoning, the tone, and whether sources or browsing indicators were shown. That last detail matters because a sourced answer comes from live retrieval and can change within weeks, while a memory-fed answer moves only when the model does. Outliers get re-run and both results kept, so the baseline reflects what a typical buyer sees.

Step 2: How do we measure share of voice against competitors?

The baseline becomes a leaderboard: you and every competitor the engines named, scored on three numbers per engine and overall. Share of voice is the share of relevant answers that name a brand at all, and it is the number most brands have never seen for their own category.

MetricWhat it answersHow we record it
Mention rateHow often you appear at all when the question is relevantNamed or absent, per question, per engine, then summed
Average rankWhen you do appear, whether you are named first or fifthPosition in the answer, first to last, averaged over the answers that name you
SentimentWhether the engine recommends you, hedges, or warns people offTaken from the engine's own wording, quoted in the evidence appendix

Three numbers beat one blended score because they fail in different ways: a brand can be named everywhere and always last, or named rarely and always first. In a 25-answer panel we ran for an ecommerce software client in August 2026, the client appeared in 0 of 25 answers. The category leader appeared in 20 and the runner-up in 18, including all 5 answers on the client's strongest question. A dated leaderboard with real names turns a vague worry into a funded project, and the same panel re-runs later to show movement.

Step 3: How do we find what AI gets wrong about you and trace it to the source?

Every claim an engine makes about you gets checked against your real facts, and every wrong claim is written down with its exact wording, the source it most likely came from, and the correction. In one delivered audit, ChatGPT quoted a client's monthly prices as annual, so every quote a buyer heard read about 25% too high, and public company databases named four different people as the company's CEO. In the ecommerce panel, four of the five cost answers quoted competitor prices word for word from those competitors' public pricing pages. One engine still recommended a vendor whose domain no longer resolves, because directories and old roundups still carried the name.

Tracing is what makes a fix possible, because correcting your own site while a stale directory keeps feeding the engine the old price changes nothing. Alongside the answers we sweep the third-party surfaces the engines in your category read: software directories, review platforms, company databases, app marketplaces where relevant, and the Reddit, Quora and forum threads that rank for your buyer questions. In the ecommerce panel we checked 18 such surfaces and the client was absent from all 18, which explained the zero more precisely than any score could.

The technical checks belong to the same step: whether AI crawlers can fetch your pages, whether your key facts render without JavaScript, and whether structured data and sitemaps are in order. Our own self-audit in August 2026 found the hosting CDN returning HTTP 429 to GPTBot and Perplexity's crawlers on several pages, and homepage counters rendered as "0" in the raw HTML, so a crawler could quote zeros; both were fixed and re-tested.

Step 4: How does a finding become a fix card?

Every finding ends in a fix card, and the cards are ordered by impact so the first one you open is the one worth doing first. A card has a fixed anatomy, so a colleague who never saw the audit can execute it from the card alone:

  1. One owner role (marketing, developer, founder), so the card lands on one desk.
  2. An effort estimate, so the quick wins are visible at a glance.
  3. Numbered steps that name where each action happens: the page, the profile, the thread.
  4. The ready-to-use asset: the corrected page opening, the schema block, the correction email to the directory, the disclosed reply to the forum thread, already written.
  5. A "done when" check you can verify yourself, such as the retired pricing page returning a redirect or the directory profile showing the current plan.

A Full Audit carries at least 15 such cards, grouped in a program view into this week, this month and this quarter. Two bounds hold on every card. Outreach and community replies are written to be posted under your real name with disclosure, and no card asks you to fake a review, a date or an affiliation, because engines and communities both punish it and it would poison the re-measures.

Step 5: What does the QA gate check before delivery?

Before a report leaves, it passes a written QA checklist. Every finding must point to evidence in the appendix, every fix card must carry an owner role, an effort and a done-when check, and every claim about your company must match the ground truth you confirmed at the start. The test for each card is whether a colleague who never saw the audit could execute it from the card alone, and a card that fails gets rewritten. Laurynas Leskauskas runs every MentionShare audit and signs the QA, and the cover of every report says so.

Step 6: How do the monthly re-measures work?

After a Full Audit, the same 30 questions run again every month for seven months, on the same engines with the same recording method, and you receive a short written update with the deltas and the next actions. Re-measuring matters because the answers keep moving after the report: models update, competitors publish, and fixes land at different speeds, with a correction on a live-retrieved page showing up within days to weeks while a corrected company database propagates more slowly. Same questions, same method, every month is what makes the trend line trustworthy, and it is how you catch the month a model update quietly drops you or invents a new error.

How can you run a rough version yourself?

You can run the honest core of this method in one engine in about 15 minutes, and the table you fill in becomes the baseline every later change gets measured against:

  1. Write five buyer questions using the formula "best [your category] for [a buyer situation you serve]", plus one price question and one "[market leader] alternatives" question. Done when each reads like something a buyer would type, with none of your brand names in it.
  2. Run them in ChatGPT, one question per new chat, recording the brands named in order, whether you appeared, and whether sources were shown. Done when the five rows of your table are filled.
  3. Ask about yourself directly ("What is [your brand]?", "Is [your brand] a good choice?") and check the description, any prices quoted, and the company facts against reality. Done when every wrong claim is written down verbatim.
  4. Run three cause checks if you were absent: search site:yourdomain.com on Bing and on Google to see how much of your site each index holds; search your category's directories and review platforms for your brand; ask the engine "what is [your homepage URL] about?" to see whether the page lifts cleanly. Done when you know which of the three is the weak point.
  5. Repeat monthly with the same table. Single-run swings are normal and three months of the same absence is a signal. Done when you have two dated rows per question.

The rough version tells you where you stand in one engine on five questions. It cannot verify claims at scale, trace each error to the surface feeding it across five engines, or write the fixes, and those three jobs are where the paid audit spends its hours. What to do about each gap it reveals is covered in our guide to getting cited by ChatGPT and Perplexity.

What is in the report, section by section?

The report has six sections, in the order a busy reader needs them, and every finding carries an evidence reference that resolves to the appendix, so nothing rests on our say-so. You can read the full 25-page sample report before paying anything; its company is fictional so nothing needs redacting, and the method, scoring and assets are exactly what you receive about your own brand.

SectionWhat it tells you
Executive summaryYour overall AI visibility score, the three most damaging findings, and your first three moves, on one page
Score at a glanceNine scored components, from AI share of voice to crawler access and structured data
Engine ranking snapshotQuestion by question, engine by engine: who gets named, at what rank, and where you are absent
Program viewEvery fix on one page, grouped into this week, this month and this quarter, with owner and effort per fix
Fix planThe prioritized fix cards, each with numbered steps and its ready-to-use asset
Evidence appendixThe leaderboard, the documented AI errors with traced sources, and the technical findings

What evidence is this method built on?

The method is hands-on, and every example on this page comes from audits MentionShare delivered in 2026 for two clients, one of them an ecommerce software company, plus our own self-audit of mentionshare.com. Client names and identifying details stay out because the reports are confidential, which is also why the public sample uses a fictional company. The dated panel, the per-engine leaderboard, the traced errors and the fix-card anatomy are the same in every audit we run.

Three limitations belong next to that evidence. A panel of 5 to 30 questions on five engines samples what buyers see; it cannot list every answer they could get. Single answers vary between runs, which is why outliers are re-run and why one screenshot is never a finding. And a re-measure shows movement over months, because months is the timescale on which sources get corrected and models update.

What the audit does not do

The audit does not guarantee rankings, mentions or citations, because nobody controls what a model says and a provider who promises that is promising something they cannot deliver. It does not implement the fixes for you: the cards are written so your own team can ship them, and hands-on implementation is a separate Partner arrangement. It does not run as an always-on dashboard; it is a deep, dated photograph with monthly re-measures, and once your fixes have landed a monitoring tool is the right instrument for watching the numbers daily. And it takes days rather than minutes, because a person reads every answer, which is also why the findings hold up when your board asks where the numbers came from.

What you receive and what it costs

Both plans run the full method described above; they differ in the size of the question panel and what follows delivery. The Snapshot costs 149 EUR one-off and runs your 5 most important buyer questions across all five engines in about 3 business days, with a mini leaderboard and the errors and fixes those questions surface. Book a Full Audit within 30 days and the 149 EUR is credited in full. The Full Audit costs 690 EUR one-off and runs 30 questions. It delivers the full leaderboard, every documented error traced to its source and 15+ fix cards in 5 to 7 business days, then re-measures all 30 questions monthly for 7 months; afterwards monitoring continues at 149 EUR per month, cancel anytime.

Prices are ex-VAT, and both plans carry a 14-day money-back guarantee: if a Full Audit does not surface something you can act on, we refund it. To see the deliverable first, read the sample report; to talk it through, tell us what you sell and we reply within one business day with a scope and a price.

Why the methodology is built this way

MentionShare's AI visibility audit methodology is built around one constraint: the answers buyers get from AI are short, confident and unlogged. The only honest way to know what they say is to ask the real questions, read every answer, check every claim and write down where each one came from. A score generated in seconds cannot tell you that an engine is quoting a retired price or naming the wrong CEO, and those are the findings that change deals. Dated panels make the work repeatable, traced sources make the fixes real, and the fix-card format turns the report into work your team starts the same week.

Run the 15-minute version this week if you have never looked, and bring in the full method when you want the why and the what-to-change across all five engines. Either way, the first dated baseline is the document every later improvement gets measured against.

Frequently asked questions

How many buyer questions should an AI visibility audit test?
Enough to cover the buyer situations you actually win, with price, alternatives and trust questions included: 5 gives a first read and 30 gives a baseline you can re-measure. More questions add cost faster than insight, because each one runs across five engines and every answer gets read by hand. We scope above 30 only for multi-market or multi-language brands.
AI answers change between runs, so how can an audit be reliable?
By treating each answer as one measurement rather than the truth. We run a fixed question set across five engines on a logged date, re-run any outlier and keep both results, and read for the pattern across questions and engines. The monthly re-measures then use the same questions and the same method, so a trend over months is real while a single swing is noise.
Do I need specialized tools to run an AI visibility audit?
No. The core of the method is a set of buyer questions, the five consumer apps, and a recording table, which is why the 15-minute version above needs no budget. Tools help with the counting once you know your baseline; they do not check whether the claims in the answers are true, and that verification is the part a person has to do.
Why does a brand rank first on Google but get ignored by ChatGPT?
Because the two systems read different evidence. Google ranks your pages; an AI engine assembles a recommendation from the sources it retrieves and remembers, such as directories, review platforms, comparison roundups and community threads, and a brand with strong rankings can have an empty record on those surfaces. The audit's surface sweep and source tracing exist to find exactly that gap.
Can you pay to be recommended by ChatGPT or Perplexity?
Not inside the recommendation. The engines that have introduced ads present them as labelled units kept apart from the answer text, and the brands the answer itself names come from the sources the engine reads and remembers. The work that moves the answer is making those sources complete, current and consistent about you; paid directory placements are an advertising decision on their own merits, because engines read the profile content, ratings and category placement that a free, complete profile already provides.
Does the audit cover markets and languages other than English?
Yes, by scope. The method is the same in any language, because it runs the questions your buyers actually type, and we write those with you. Each extra market or language adds questions and surfaces to check, which is what moves the price on a custom scope; the Snapshot and Full Audit prices cover one market.
What happens if the audit finds nothing I can act on?
You get your money back. Both plans carry a 14-day money-back guarantee, and the Full Audit's promise is specific: if it does not surface something you can act on, we refund it. In practice a brand already named everywhere with accurate claims still gets a ranked list of accuracy and rank fixes, and that outcome is rare enough that we are glad to stand behind the guarantee.

Want to know if AI recommends you?

Get an AI-visibility audit + fix plan. See whether ChatGPT, Claude, Perplexity and Google AI Overviews name your brand, and exactly what to change. Run on our current 2026 methodology, updated as the engines change.