Resources · Guide 5 of 5

AI visibility testing.

Google Search Console tells you how you rank. Nothing tells you what ChatGPT, Claude, Gemini, Copilot, or Perplexity say about you when a buyer asks — because that conversation never touches your analytics at all. Here's how to check for yourself, on a cadence, without a vendor's dashboard.

TL;DR

There's no Search Console for the answer layer. Analytics tools measure clicks and impressions from a results page; they have no mechanism to observe what an engine actually says about you inside a generated answer, because that conversation never touches your site. The only way to know is to ask — put a small, stable set of questions to a handful of engines, record who gets named and who doesn't, and repeat it on a schedule. A single response proves nothing; engines are inconsistent by design and disagree with each other more than most people expect, so one check tells you about one engine on one day, not your visibility generally. And the finding worth prioritizing above everything else isn't silence — it's a confidently wrong answer. An engine stating the wrong price, the wrong service area, or the wrong specialty is actively harmful in a way that simply not being mentioned never is.

There's no dashboard for this

For most of the last two decades, "measuring visibility" meant checking a dashboard. Search Console reports impressions and clicks. Analytics reports sessions and referrers. Every number in every one of those tools is downstream of a results page a person actually looked at, and usually clicked through from.

Answer engines break that model completely. When someone asks ChatGPT, Claude, Gemini, Copilot, or Perplexity a question and the engine names three businesses in a synthesized answer, none of that touches your website's analytics. No impression fires. There's no click to log, or fail to log, in any tool built to watch a results page. The exchange happens entirely inside the engine's own interface — and if you weren't one of the names it used, there is no record anywhere that you were even in the running.

Even Google, which controls both the results page and the AI Overview sitting on top of it, doesn't hand you this view. AI Overview appearances are folded into your general Search totals rather than broken out on their own, so even the company running the feature doesn't give you a way to isolate it.

This is a different question from the one a readiness check answers. The visibility ecosystem covers why AuditSpark's own site can score a flawless 100/100 on readiness and still measure 17.3/100 on actual visibility — two different dials, both real. This piece is about the second one: how to read it yourself, informally, without waiting on a vendor's tooling to do it for you.

Engines disagree more than you'd expect

Check one engine and you've measured one engine, not "AI visibility" in general. A 2026 measurement study that ran 602 controlled prompts across several major platforms and analyzed more than 21,000 resulting citations found very low overlap between which sources different engines actually cited for the same questions. An engine that names you reliably might be the exception, not the floor — and an engine that never does might be the outlier in the other direction.

That same research draws a distinction worth borrowing whole. It separates citation selection — does an engine cite a given source at all — from citation absorption — once cited, how much that source actually shapes the wording of the final answer. A site can be cited constantly and contribute almost nothing to what the answer ends up saying, or get cited rarely and dominate the one answer where it does appear. A plain yes/no mention count misses this entirely. If you get past a first pass at this, it's worth noting not just whether you're named, but how much of the answer's substance actually traces back to you when you are.

Behavior isn't fixed, either. The same question, put to the same engine minutes apart, can come back with a different answer — different names, different order, sometimes a different framing entirely. Treat any one check as a snapshot that ages, not a permanent rating.

Design a small, stable probe set

The practical version of "ask the engines" is a short list of questions, put to several engines the same way every time, scored the same way every time. Fifteen to twenty-five questions is enough to start, spread across four kinds of intent:

#Question typeWhat it actually tests
1DiscoveryUnbranded category questions a buyer asks before knowing any names — "best CRM for a small law firm"
2ComparativeVersus and alternatives questions, once a buyer has a shortlist
3Recommendation-seekingAdvice questions with real constraints attached — budget, location, team size
4Branded factualDirect questions about your own business, specifically to catch misrepresentation

That fourth type is the one most people skip, and it's the one that matters most. Asking an engine what it thinks it knows about your pricing, your service area, or your specialty is the only way to catch it confidently getting one of those wrong — and a wrong answer stated with total confidence reaches a buyer as fact.

For each question, on each engine, the same handful of things is worth recording every time: whether you're mentioned at all, whether you're cited as a source or just named in passing, roughly where you land in the answer, which competitors show up instead of or alongside you, whether the framing reads positive, neutral, or negative — and, above all, whether anything the engine states about you is factually wrong. Escalate factual errors ahead of everything else on that list. A confidently wrong answer does more damage than being left out, because being left out is passive and a wrong answer is actively misleading a buyer.

Two rules make the whole exercise worth repeating instead of a one-off curiosity. Hold the exact same question set stable across rounds — changing what you ask between checks breaks the comparison, because you're no longer measuring the same thing twice. And treat any single response as an anecdote, not a trend: engines are non-deterministic, so only a pattern across multiple questions and multiple runs means anything. If something on your site changes and a later check looks different, the honest framing is "this changed," never "this is why" — testing like this can show you a shift, but it can't prove what caused it.

What this looks like in practice

Most businesses currently have zero visibility into this layer at all, so even a basic, informal version of this check is genuinely new information — not a rounding error on something you already know. A few concrete steps get you there without any tooling:

  • Write the question set this week. Fifteen to twenty-five questions, covering all four intent types above, phrased the way a real buyer would actually ask them — not the way you'd phrase them about your own business.
  • Run it through two or three engines by hand. ChatGPT, Claude, Gemini, Copilot, Perplexity — whichever few are realistic for you to check regularly. A spreadsheet with one row per question per engine is plenty of rigor for a first pass.
  • Flag factual errors separately from everything else. Wrong pricing, wrong location, wrong specialty gets fixed first, ahead of anything about tone, ranking, or how many competitors got named instead of you.
  • Repeat it on a cadence, not once. Monthly or quarterly is enough to catch a real shift. A single check tells you about a single moment; a repeatable one is what turns this into evidence instead of an anecdote.

Two habits quietly undermine this even when the intent is right. Only asking questions that already name your brand tests recognition, not discovery — a much easier thing to pass, and a different question than whether a buyer who's never heard of you would find you. And treating every mention as equally good flattens a real difference: being named third, behind two competitors, is not the same result as being named first, even though both count as "mentioned" on a simple yes/no basis.

Frequently asked

Isn't this what SEO rank tracking already does?

No. Rank tracking measures your position in a results list a search engine returns. There's no list here — the engine writes a synthesized answer and either names you inside it or doesn't. A rank tracker has no visibility into that text at all; it wasn't built to read it.

How many engines and questions do I actually need to check?

Fewer than feels rigorous. A handful of engines and fifteen to twenty-five stable questions spanning discovery, comparative, recommendation, and branded-factual intents is enough to see real patterns. Asking more questions matters less than asking the same ones consistently over time.

Can this prove that something I changed improved my visibility?

No, and it shouldn't claim to. Engines are non-deterministic and influenced by far more than any one thing you control. This kind of testing can show you that something changed between two checks; it can't establish that your change was the cause. Keep the framing to "this changed," not "this is why."

See your own numbers.

The free readiness check reads your site the way an AI engine would. Four to five minutes, no charge, one concrete finding.