How To Check A Page Is In AI Retrieval (Snippet Test)

· The Cresia team

Yes, you can check whether an AI system can retrieve one of your pages without any vendor dashboard. Paste a distinctive passage from the page into a chatbot that can search the web, ask it to find exact matches, and see whether your URL comes back. Search Engine Journal's Chris Green describes this snippet test as a way to separate retrieval problems from everything else: if your URL returns, retrieval isn't what's holding the page back.

Key takeaways

  • A snippet test tells you whether a page can be retrieved. It does not tell you whether the page will be cited in an answer.
  • If your URL comes back, stop debugging access and work on selection: clear passages, specific claims, consistent entities.
  • If it doesn't, check crawler access, rendering, indexing and canonicals before you rewrite anything.
  • Use several distinctive passages, several tools and several dates, and log each run like any other experiment.
  • One chatbot answer is a probe, not a ranking report. Outputs vary between runs, so look for patterns.

What does the snippet test actually check?

It checks one narrow thing: when a system that searches the web is handed a string that only your page contains, does it find that page? That's a retrieval question. Nothing more.

The reason it's useful is the gap it fills. Google Search Console shows you impressions and clicks for Google's results. There's no equivalent report from most AI assistants that says here are the queries where we looked at your page. Teams end up guessing. Someone asks a chatbot a question in their category, doesn't see the company named, and concludes the content needs a rewrite. Maybe. Or maybe the crawler is blocked, the page renders its main copy with JavaScript, or the URL the system knows about is a parameterised duplicate.

The snippet test cuts through that. You pick a sentence that's unusual enough that no other page would plausibly contain it: a specific figure from your own data, a named framework you coined, a slightly odd phrasing in a product description. Then you ask for exact matches. Two outcomes matter. Your URL appears, so the page is in the pipeline. Or it doesn't, so you've got a concrete, checkable problem upstream of content quality.

What it doesn't check is preference. A page can be perfectly retrievable and still never appear in an answer to a real question, because plenty of other pages were retrieved too and something else was chosen.

Why retrieval and citation are different problems

Think of three stages. First, the system has to be able to fetch or index your page. Second, when a user asks something, the system has to pull a set of candidate pages and passages. Third, it has to decide which of those candidates to lean on and whether to name them.

The snippet test sits at stage one and partly stage two. It uses a query so specific that your page should be the only candidate. That's the point. Real questions are broader, so your page competes with a dozen others.

This matters for how you spend time. A content team that treats every missed mention as a writing problem will burn weeks rewriting pages that were never fetched. A technical team that treats every missed mention as an access problem will keep tuning robots rules on pages that are retrieved fine and simply say nothing quotable.

Keep the split in your head and in your tickets. Retrieval failures belong to whoever owns the site's infrastructure and templates. Selection failures belong to whoever owns the content and the brand's entity footprint. Mixing them creates arguments with no owner.

The stages also fail differently. Retrieval failure is close to binary and stable: the page is either reachable or it isn't, and it tends to stay that way until something changes. Selection is probabilistic and shifts with the question wording, the tool and the day. Expect the first to be diagnosable in an afternoon and the second to need a running programme.

How do you run a snippet test properly?

A single paste into a single chatbot is a start, not a test. Here's a version that gives you something you can trust and repeat.

  1. Choose passages that only you could have written. A sentence containing your own survey figure, an internal method name, or a precise product specification works. Generic statements like 'AEO is growing in importance' will match a hundred pages and prove nothing.
  2. Pick passages from different parts of the page. A sentence from the intro, one from the middle and one from a table or list. Some systems trim or ignore parts of a page, and you'll see it.
  3. Use more than one tool. Try at least two search-enabled assistants, because they don't share an index or a crawler. A pass in one and a fail in another is a real finding.
  4. Ask for exact matches, and ask for the URL. Phrase it as a request to find pages containing that exact text and to list the sources. Be wary of an answer that paraphrases your passage without a link; that can be the model's memory, not retrieval.
  5. Repeat on a different day. One run can be a fluke in either direction. Two or three runs across a week tell you whether the result is stable.
  6. Record everything. Passage, tool, date, whether the URL appeared, and which URL variant appeared.

The last point catches a common surprise. The URL that comes back may be a tracking-parameter version, an old path, or a syndicated copy on another site. That's a canonicalisation problem, and it's worth knowing about even when the test technically passes.

One caution on interpreting a pass. If the tool returns your passage from a cached or older version of the page, retrieval is working but freshness may not be. Change a distinctive sentence, wait, and test the new wording to see how quickly updates are picked up.

How do AI answer engines choose which sources to cite?

No vendor publishes a full recipe, and the details change, so hold any claim about exact weighting loosely. What can be said from how these systems are generally built is this: the assistant turns your question into one or more searches, gets back a candidate set of pages, reads passages from them, and composes an answer that may or may not credit the sources it used.

A few consequences follow.

Passages matter more than pages. The system often pulls a chunk of text that answers the question, not the whole document. A page whose useful sentence is buried in a rambling paragraph gives the system less to lift than one where a heading asks a question and the next two sentences answer it.

Agreement matters. When several retrieved sources say the same thing, that claim is easier to state confidently. If your page makes a claim nobody else does and offers no support, it's harder to use, even if it's true.

Specificity helps. Named entities, concrete numbers with a clear origin, and definitions that don't depend on surrounding context are easier to quote on their own.

Order is less decisive than people assume. Search Engine Journal reported an AI citation test under the headline that source order matters less than it looks. Read the full write-up for the method and scope before leaning on it, but the direction fits the passage-first picture: being the first result isn't the same as being the source that gets used.

The practical read: retrieval gets you into the room. What you say, and how cleanly you say it, decides whether you're quoted.

What should you fix when your URL doesn't come back?

Work from the cheapest, most mechanical check to the most editorial. Most failures live in the first half of this list.

  • Crawler access. Check your robots.txt for rules that block the user agents of the AI systems you care about, and confirm the current names in each vendor's own documentation, since they change. Also check what your CDN or bot-management layer does. Search Engine Journal covered Cloudflare's move to write robots.txt for site owners, which is a reminder that access rules may be set somewhere other than the file you're looking at.
  • Rendered content. View the page with JavaScript off. If the passage you tested isn't in the initial HTML, a crawler that doesn't execute scripts won't see it.
  • Indexing and canonicals. Make sure the page is indexable in search engines, returns a clean 200, and points its canonical at itself. Assistants that lean on a search index inherit its problems.
  • Blocking by status. Rate limits, geo rules and challenge pages can return 403 or 429 to automated visitors while looking fine to you in a browser. Look at server logs for those user agents.
  • Duplicates and syndication. If a partner site republishes your article, the copy may be what gets retrieved. Use canonical tags and, where you can, ask for attribution.
  • Thin or gated pages. Content behind a login or a heavy interstitial isn't retrievable in any useful sense.

Only after those pass should you look at the writing. Then the moves are editorial and modest: put the direct answer near the top of the section, define terms in plain sentences, and give the numbers you cite a clear origin.

Growth teams juggling technical debt and content calendars at once often need a single owner for this list. The growth teams solution page is a reasonable place to see how that kind of cross-functional work is usually organised.

How do you measure AI retrieval over time?

A snippet test is a spot check. To make it a measurement, turn it into a small recurring routine with fixed inputs and a log.

Check What you record Cadence Owner
Snippet test on key pages Passage, tool, date, URL returned or not, URL variant Monthly, and after any template or CDN change SEO lead
Crawler access review Robots rules, CDN bot settings, log evidence of AI user agents Quarterly, and after infrastructure changes Web engineering
Prompt panel A fixed set of real customer questions, which brands and URLs appear Fortnightly Content or AEO lead
Referral traffic Sessions arriving from AI assistants, tagged consistently Continuous, reviewed monthly Analytics lead
Freshness lag Days between editing a distinctive sentence and seeing it returned After major page updates SEO lead

Two notes on this. The prompt panel is where selection shows up, so keep the questions stable and phrase them the way customers do, not the way your product marketing does. And referral tagging only works if the tracking is consistent; a written tracking specification keeps the channel definitions from drifting between teams.

If you want visibility work handled in one place, MediaPilot is the Cresia product built around media and AI-search visibility; the product page describes what it covers.

Avoid reading trend lines out of small samples. Five prompts run once a month will swing for reasons unrelated to your work. Widen the panel and the time window before you report a movement to leadership.

What should you avoid doing?

A few habits cause more wasted effort than any missing tactic.

Don't treat the test as a score. Passing tells you the page is reachable. Reporting it upward as 'we're visible in AI search' overstates it, and the first missed mention will damage trust in the whole programme.

Don't pick a snippet that isn't distinctive. If the sentence could appear on competitors' pages, a returned URL might be theirs, and a missing one might just be crowded out.

Don't rewrite copy first. It feels productive and it's the most visible change you can make, which is exactly why it gets done before the boring access checks. Check the boring things first.

Don't open every door at once. Letting all crawlers in because a test failed can be the wrong call if you have licensing or load concerns. Decide by crawler and by purpose, and keep a record of why.

Don't chase one assistant's quirks. Behaviour differs by tool and changes without notice. Fixes that improve access, clarity and consistency help across the board. Fixes that target one system's observed habit tend to expire.

And don't pay for tooling to answer a question you can answer in ten minutes. The snippet test costs nothing. Buy software when you've outgrown the manual routine, not before.

Frequently asked questions

Does a URL coming back mean I'll be cited in answers?

No. It means the page can be found when the query is specific enough to match only it. Real questions bring in competing pages, and the system then chooses among them. Retrieval is necessary for citation, not sufficient.

Which chatbot should I use for the test?

Use at least two that can search the web, since they rely on different indexes and crawlers. A tool answering only from its training data can't confirm retrieval, so make sure browsing or search is switched on. Check each vendor's documentation for how search works in that product.

What if the tool returns my passage but names a different URL?

That points to duplication or syndication. Another page carries your text, or a variant of your own URL is what got indexed. Fix canonicals, check partner republishing, and rerun the test after the changes have had time to be picked up.

How often should I rerun the test?

Monthly for your most important pages is a sensible default, plus after any change to templates, rendering, robots rules or your CDN's bot settings. Those are the changes most likely to break retrieval without anyone noticing.

Is this a replacement for tracking AI referral traffic?

No. The snippet test checks access. Referral analytics show whether people arrive from assistants, and a prompt panel shows whether you're being named. You need all three, because each answers a different question.

Sources

  • https://www.searchenginejournal.com/checking-a-page-is-part-of-a-retrieval-pipeline-for-ai/589284/
  • https://www.searchenginejournal.com/ai-citation-test-finds-source-order-matters-less-than-it-looks/589806/
  • https://www.searchenginejournal.com/cloudflare-will-write-your-robots-txt-and-it-has-a-point/589262/

See Cresia on your own use case.

30 minutes, your team, your questions.

Request a demo