How to Measure AI Overviews and LLM Visibility

· The Cresia team

Enterprise teams are measuring AI Overviews and LLM visibility with three imperfect tools at once: tracked prompt sets, referral analytics and manual citation audits. None is complete alone, so the workable approach is a small, stable set of prompts watched as a trend, not a rank. Search Engine Journal is promoting a webinar where Tom Capper of STAT covers the main measurement problems, which is a fair sign the question is still open.

Key takeaways

  • Track a small, fixed prompt set as a trend line. Single-run answers are too variable to treat as a rank.
  • Combine three signals: sampled prompt tracking, referral and branded-search analytics, and manual citation audits.
  • Answer engines cite passages that stand alone, so write sections that answer one question completely.
  • Layout tests, like the citation card placement Google keeps trying, mean any fixed click-through assumption will go stale.
  • Assign owners before buying a tool: content, SEO, analytics and brand each hold part of the measurement.
  • Skip vanity scores and one-off screenshots. Neither survives a second run of the same prompt.

What the AI search measurement conversation looks like right now

Search Engine Journal, in a post by Loren Baker, listed a webinar in which Tom Capper of STAT walks through the biggest AI search measurement challenges and a tactical fix for each. It runs October 14 at 2 p.m. ET. That's a session announcement, not a dataset, so there's no new finding to react to here. What matters is the framing: measurement has become the stated bottleneck for enterprise SEO teams, ahead of tactics.

Two other items from the same news cycle show why. Search Engine Roundtable reported that Google is again testing citation cards at the bottom of AI Overviews instead of on the right side. Search Engine Journal reported that CallRail is making ChatGPT ads measurable for SMBs and agencies. One story is about where your link appears. The other is about vendors racing to hand marketers a number.

Put those together and the picture is plain. The surfaces keep moving, and the tooling to count what happens on them is being built while you use it. A team that waits for stable measurement will wait a long time. A team that builds a modest, honest measurement habit now will have a baseline when the tools settle.

Why AI search measurement breaks rank-tracking habits

Classic rank tracking works because a query returns a list, and your URL sits at a position. Both assumptions weaken in AI search. The same prompt can produce different answers on different runs, for different users, in different sessions. Sometimes your brand is named without a link. Sometimes a link appears without your brand being described accurately. Sometimes neither happens.

That creates three specific problems for a reporting deck.

First, position means little. A citation in an AI Overview isn't rank three or rank seven. It's a presence or absence, and where it sits on the screen is a layout choice Google can change, as the citation card tests show.

Second, sampling is unavoidable. You can't observe every prompt a customer types. Any tracker is showing you a sample of prompts, run at some frequency, from some location, on some model version. Ask the vendor how each of those is set, and read the answer in their own documentation.

Third, the outcome you care about often happens off your site. A buyer reads a summary that names three vendors, then searches for one of them by name. Your analytics logs a branded search or a direct visit, not an AI referral. If you only count sessions tagged to an AI source, you'll undercount the influence and overcount the noise.

The honest conclusion is that AI visibility is a probabilistic measure. You're estimating a rate: out of a fixed set of prompts, how often does an engine mention or cite you, and how does that move after you change something? That's a fine thing to report. It just isn't a rank.

How AI answer engines select and cite sources

The mechanics differ by product, and they change, so check each vendor's own documentation for what it says about crawling and sourcing. The general pattern is stable enough to plan around, though.

Most answer engines that show citations do some form of retrieval. They take the prompt, look for candidate pages in a search index or their own crawl, pull out relevant passages, and have a model write an answer grounded in those passages. Some answers come purely from what the model already learned in training, and those often cite nothing at all.

That has consequences for how you write:

  • Passages get lifted, not pages. A section that answers one question fully, with its subject named in the first sentence, is easier to quote than a paragraph that depends on the three above it.
  • Being retrievable comes first. If a crawler can't fetch the page, or the content only appears after heavy client-side rendering, it can't be cited. Check robots rules and rendering before you touch copy.
  • Clarity about entities helps. If your page says exactly which product, region and use case it covers, the model has less to guess.
  • Agreement across sources matters. Engines tend to be more comfortable stating a claim that several independent pages support. A fact that only your own site asserts is a weaker candidate.

None of this is a trick. It's the same discipline as writing a good reference page, applied with the knowledge that a machine will quote you out of context. Write every section as if it will be read alone, because it might be.

How to measure AI Overviews and LLM visibility

Build the measurement in layers. Each layer is imperfect, and the value comes from seeing where they agree.

  1. Define the prompt set. Pick 30 to 100 prompts that match real buying and research questions, split into branded, category and comparison prompts. Freeze the list. Change it on a schedule, not on a whim, or your trend line means nothing.
  2. Run it repeatedly. Run each prompt several times per period, across the engines your buyers use. Record whether you're mentioned, whether you're cited with a link, which competitors appear, and whether the description is accurate.
  3. Watch the referral and demand signals. Tag AI referral sources where they exist, and watch branded search and direct traffic for lift after content changes.
  4. Audit citations by hand each month. Read the answers. A tracker can tell you that you were cited. It can't tell you the summary called your product something it isn't.

Here's how the layers compare:

Signal What it tells you Main weakness Cadence
Tracked prompt set Mention and citation rate over time Sampled, sensitive to model changes Weekly
AI referral traffic Clicks that arrive from answer engines Misses influence without a click Weekly
Branded search and direct Downstream demand lift Slow, mixed with other campaigns Monthly
Manual citation audit Accuracy and tone of descriptions Small sample, labour-heavy Monthly
Competitor share of answers Who fills the slots you don't Depends on the fixed prompt set Monthly

One more rule: report ranges and direction, not decimals. If your mention rate moved from roughly one in four runs to roughly one in three across a stable set, say that. Don't publish a figure to one decimal place from a sample that can't support it.

Teams that already maintain clean analytics have an edge here, because AI referral tagging and event naming depend on a tidy foundation. If yours is patchy, start with a tracking specification before you add another dashboard.

What to change on your pages and in your workflow

Measurement without changes is just a hobby. These are the moves that tend to be worth making, roughly in order of effort.

  1. Fix access first. Confirm which crawlers you allow, that key pages return real HTML, and that canonical tags point where you think they do.
  2. Rewrite the top of your best pages. Lead each section with the direct answer, then the nuance. If the heading is a question, the next sentence should answer it.
  3. Make comparison content honest and specific. Engines get asked for shortlists, as Search Engine Land's coverage of ChatGPT shortlists suggests. A page that states who a product is for, and who it's not for, is easier to include than one that claims to suit everyone.
  4. Keep facts consistent everywhere. Product names, pricing model, integrations and locations should match across your site, documentation, profiles and press. Inconsistency is how wrong summaries happen.
  5. Add real evidence. Original examples, named use cases and first-hand detail give a model something worth citing that a competitor's generic page can't offer.
  6. Test one change at a time. If you rewrite thirty pages and the mention rate moves, you've learned nothing you can repeat. Change a cluster, hold another as a control, and compare.

That last point is where experimentation discipline pays off. The logic of A/B testing applies even when the outcome is a citation rate: define the change, define the comparison, and wait long enough to see past run-to-run noise.

Who owns which part of AEO measurement

AI search visibility doesn't sit neatly inside one team. Content writes the passages, SEO owns access and structure, analytics owns tagging and reporting, and brand owns whether the descriptions are right. When nobody owns the whole, the prompt set drifts and the report gets ignored.

Task Suggested owner Output Review point
Prompt set definition SEO lead Frozen list of prompts by intent Quarterly
Crawler and rendering checks Technical SEO Access report After each release
Passage rewrites Content lead Updated pages with change log Per batch
Referral and event tagging Analytics lead Clean AI source reporting Monthly
Accuracy audit of answers Brand or product marketing List of wrong or stale claims Monthly

If this sounds like a marketing operations problem, that's because it is one. Someone has to keep the prompt list, the tagging and the change log in sync across teams. The marketing operations primer covers that coordination role, and teams that want tooling for AI-search visibility can look at MediaPilot to see whether it fits.

What not to do when optimising for AI search

A few habits will cost you time and credibility.

  • Don't screenshot one answer and call it a result. Run the prompt again and it may look different. A single capture is an anecdote.
  • Don't chase a composite visibility score you can't explain. If nobody can say what goes into it, nobody can act on it or defend it to finance.
  • Don't rewrite your whole site for a layout that may change. Google's citation card testing shows placement is in flux. Write for clear passages, not for a card position.
  • Don't stuff pages with question-and-answer blocks that repeat each other. Thin duplication reads as filler to people and to models.
  • Don't treat AI referral traffic as the whole story. Many influenced buyers arrive through a branded search later.
  • Don't buy a tool before you've written the prompt set. The tool will happily measure the wrong questions at scale.

And a note on paid surfaces. As ChatGPT ads become measurable, as Search Engine Journal reports for CallRail's SMB and agency offering, keep paid and organic answer visibility in separate reports. They answer different questions, and blending them hides both.

Frequently asked questions

Can you rank track AI Overviews the way you track blue links?

Not in a meaningful way. An AI Overview is a generated answer with citations, and its content and layout vary between runs and over time. Track presence, citation and description accuracy over a fixed prompt set instead.

How many prompts should a fixed prompt set include?

Enough to cover your main buying questions across branded, category and comparison intent, and few enough that you can review the answers by hand. A few dozen is a reasonable start for most teams. Expand only when you can maintain it.

Do AI referrals show up cleanly in analytics?

Only partly. Some engines pass identifiable referrers and some clicks arrive as direct traffic or later branded searches. Check how each platform documents its referral behaviour, and read the referral count as a floor on influence, not the total.

Is GEO different from traditional SEO?

The foundations overlap heavily: access, clear content, consistent facts and credible sources. GEO adds attention to how passages read when quoted alone and to how you're described in generated answers. Treat it as an extension of good SEO with a different measurement problem.

Sources

  • https://www.searchenginejournal.com/how-are-enterprise-seo-pros-measuring-ai-overviews-llms-webinar/590971/
  • https://www.seroundtable.com/google-ai-overviews-citations-cards-bottom-42177.html
  • https://www.searchenginejournal.com/callrail-makes-chatgpt-ads-measurable-for-smbs-and-marketing-agencies-spn/589843/

See Cresia on your own use case.

30 minutes, your team, your questions.

Request a demo