How to Measure AI Search ROI Without Perfect Data
You can measure AI search ROI, but not as a single number. The workable approach is a chain of weaker signals (visibility in answers, referral visits, branded demand and what buyers tell sales) judged by whether they rise and fall together over a quarter.
Key takeaways
- No tool gives you clean AI search attribution today. A set of signals that move together is the honest substitute.
- Ahrefs Blog reports Forrester's finding that 94% of business buyers use AI search when buying. That supports spending on measurement; it doesn't support revenue forecasts.
- Track a fixed prompt set, AI referral visits, branded search demand and what buyers say on sales calls. Each covers a gap the others leave.
- Answer engines lift passages that state one thing plainly and are backed by mentions elsewhere, so page structure and off-site presence both count.
- Skip vanity scores, single-prompt rank tracking and any report that claims exact attribution.
- Fix your tracking plan before you build dashboards. Inconsistent event names and UTM rules poison every AI number downstream.
What the new buyer-behaviour numbers actually change
Ahrefs Blog, in its piece on measuring AI search ROI, cites Forrester's finding that 94% of business buyers now use AI search when buying. According to that report, they rate it above vendor websites, product experts and sales reps. On the consumer side, Ahrefs Blog cites NielsenIQ putting the share of consumers who research products with AI at 42%.
Those are survey figures about behaviour. They say your audience is already asking assistants for recommendations. They don't say how much of your pipeline passes through an AI answer, and no one outside your own data can tell you that.
So what changes? The internal argument. A year ago the question in budget meetings was whether AI search deserved attention at all. Now it's closer to this: what will you count, and what will you accept as evidence? The numbers above are a good reason to fund measurement. They're a poor reason to promise a revenue figure.
Here's a situation many teams recognise. A head of marketing at a mid-size software company is asked by the CFO what AI search is worth. Saying 94% of buyers use it is true and useless, because it doesn't touch the company's own funnel. Saying we appeared in a certain share of a fixed set of buyer questions this month, and our branded search demand and demo requests that mention an assistant moved in the same direction, is smaller and far more defensible.
Why AI search ROI is hard to attribute
Traditional search attribution works because the click is the event. A person searches, clicks, lands, converts. AI answers break that chain in several places.
- The answer often ends the research step. A buyer reads a shortlist, doesn't click, and types your brand name into a browser two days later. That visit shows up as direct or branded organic.
- Referrer data is inconsistent. Some assistants send a referrer, some strip it, and in-app browsers behave differently again. Check each vendor's documentation and your own analytics rather than assuming.
- Answers vary. The same question can produce different recommendations by day, by phrasing, by location and by user history. One screenshot proves nothing.
- Several touches are blended. A buyer might see you in an answer, read a review site, then request a demo from a paid ad. Credit goes to the ad.
None of this is a reason to give up. It's a reason to stop expecting a last-click number and to build a measurement design that admits where the gaps are. Marketing operations teams will know the pattern: when the data can't tell you the answer directly, you agree the definitions up front and hold them steady. The marketing operations primer covers that discipline in general terms.
How AI answer engines choose and cite sources
The mechanics differ by product, and they change often, so treat this as a working model rather than a spec.
An assistant can answer from what it absorbed in training, or it can run live searches and compose an answer from the pages it retrieves. Most commercial questions (best tools for X, which vendor suits Y) lean on retrieval, because buyers want current information. In retrieval mode the engine fetches candidate pages, reads passages, and builds a response with some of those pages cited.
Three things follow from that.
Passages get lifted, not pages. A paragraph that names the product, states what it does, who it's for and what it costs, in plain sentences, is easy to quote. A paragraph of positioning language isn't. Write sections that answer one question completely.
Third-party presence matters. Engines tend to cross-check. If review sites, forums, analyst notes and press describe you consistently, you're easier to recommend. If the only description of you is your own homepage, you're easier to skip. Search Engine Land has reported that ChatGPT and Google's AI favour large retailers in shopping answers, which is a reminder that established brands carry weight. A smaller brand shouldn't try to out-argue them on broad terms. Narrow, specific comparisons and use cases are where there's room.
Access is part of the picture. If a crawler can't read your pages, it can't cite them. Search Engine Land has also covered publishers pushing Congress over AI crawlers that hide their identity, so crawler behaviour is contested ground. Check your server logs and each vendor's documentation for the user agents they say they use, and decide deliberately which to allow in robots.txt.
A measurement plan you can run this quarter
The aim is four layers, each imperfect, each reported with its limits stated. Search Engine Land's coverage of benchmarking visibility in ChatGPT shortlists points the same way: pick a stable set of questions and watch it over time instead of reacting to single answers.
| Layer | Question it answers | Signal | Suggested owner |
|---|---|---|---|
| Visibility | Do we appear in answers to buyer questions? | Share of a fixed prompt set where the brand is named or cited, re-run on a schedule | SEO / AEO lead |
| Referral | Do any answers send visits? | Sessions from assistant referrers, landing pages, engagement | Analytics lead |
| Demand | Is brand interest rising? | Branded search impressions, direct visits, new-vs-returning mix | Growth lead |
| Pipeline | Do buyers say AI found us? | Self-reported source on forms, sales call notes, deal context | Sales ops / RevOps |
A few notes on running it.
- Build the prompt set once. Write thirty to fifty questions a real buyer would ask, grouped by stage: problem, category, comparison, vendor shortlist. Freeze the wording. Add questions in a new version rather than editing old ones.
- Run it on a schedule, several times. Because answers vary, run each prompt more than once per cycle and report a range or an average, not one result.
- Log who else appears. The competitors named next to you tell you more than your own mention does.
- Add a free-text field. On demo forms, ask how the buyer first heard of you and let them type. Read the answers monthly. Assistants will show up in the wording before they show up in your referrer reports.
- Tag the sales side. Ask reps to note when a prospect says they asked an assistant. A simple field beats none.
Then compare the layers. If visibility rises, branded demand rises and self-reported mentions rise, you have a case. If only visibility rises, you may be winning answers that nobody acts on. If demand rises without visibility, something else is moving the needle and you shouldn't credit AI search.
Tools exist to automate the visibility layer, and MediaPilot is Cresia's product for media and AI-search visibility. Whatever you use, check the vendor's own documentation for how it samples answers, because that method decides what your numbers mean.
What to change on your pages and in your stack
Measurement tells you where you stand. These changes move it.
- Rewrite your core pages as answers. For each product or service page, open with two plain sentences saying what it is, who it suits and what sets it apart. Put pricing basis, limits and integrations in text, not just in images or tabs.
- Make comparison content honest. Buyers ask assistants for alternatives. A fair comparison page that admits where a competitor is stronger is more quotable, and more trusted, than a one-sided one.
- Give every fact one canonical home. If three pages describe your pricing three ways, an engine may pick any of them. Choose the source of truth and point the others at it.
- Earn mentions where buyers look. Review sites, community threads, industry roundups. This is slow and it can't be faked, but it's the layer most teams neglect.
- Keep pages fresh where facts change. Out-of-date specs and old prices get repeated.
- Tidy the tracking underneath. If your analytics tool names events differently on different templates, or UTM parameters are applied by habit rather than rule, your referral layer will be noise. A written tracking specification fixes that, and OmniSpec is Cresia's product for analytics governance.
Growth teams that own both content and reporting often have the clearest view of how these pieces connect; the growth teams page outlines how Cresia approaches that work.
What to ignore and what not to do
The field is young, vendors are loud, and some of the advice does harm.
- Don't chase a single AI visibility score. Composite scores hide their weights. Report the raw share of your prompt set and the competitor list instead.
- Don't track one prompt daily. Answers fluctuate. A day-to-day wobble on one question is not a trend.
- Don't claim exact attribution. If a report says AI search drove a precise revenue figure, ask how. If the method is a last-click rule on a small referrer, the figure is a floor, not a total.
- Don't stuff pages with question-and-answer blocks. Writing for passages is not the same as spamming FAQs. Each section should earn its place.
- Don't assume the old playbook is dead. Search Engine Land has run a piece debunking AI search myths; the broad lesson holds that crawlable, useful, well-structured pages remain the base. AI search sits on top of that work, not instead of it.
- Don't block crawlers by default and then wonder why you're absent. Make the access decision on purpose, document it, and review it when vendor policies change.
- Don't over-index on today's numbers. Product behaviour and citation patterns are shifting. Hold your methodology steady, and expect to revise conclusions.
Reporting AI search results to finance and leadership
A short report works better than a dashboard of twenty tiles. One page, four parts: what we measured and how, what moved, what we think it means, and what we can't know.
The last part matters more than it looks. Saying that referral traffic undercounts AI influence, and why, builds trust. A leadership team that has been told the limits will not be surprised when a number doesn't reconcile.
Tie each layer to a decision. If visibility is low on comparison prompts, the decision is to fund comparison content. If branded demand is rising but referral is flat, the decision is to add the self-reported source field and read it. If sales hear about assistants but your prompt set shows nothing, your prompts aren't the questions buyers ask, so rewrite them from call transcripts.
Creative and media teams get pulled in here too, since the same messages show up in ads, landing pages and answers. The media teams page is a starting point if that overlap is yours to manage.
Frequently asked questions
Can you really measure ROI from AI search?
Not to the decimal. You can build a defensible case from several signals that move together: visibility in a fixed prompt set, referral visits, branded demand and what buyers report. Treat the result as evidence for a decision, not as an accounting figure.
Do the Forrester and NielsenIQ figures tell me what AI search is worth to my business?
No. Ahrefs Blog reports Forrester's 94% for business buyers and NielsenIQ's 42% for consumers, and both describe how people research. Your own funnel data decides what that behaviour is worth to you.
How often should I re-run a prompt set?
Often enough to see a trend and rarely enough to avoid reacting to noise. Many teams settle on a weekly or fortnightly cycle, with each prompt run more than once. Keep the wording fixed so changes reflect your visibility, not your phrasing.
Should AEO and GEO have a separate budget from SEO?
Usually the work overlaps heavily, because the same pages, structure and off-site mentions serve both. Separate the measurement, so you can see what AI search is doing, but be wary of splitting the team in a way that creates two versions of every page.
Sources
- https://ahrefs.com/blog/ai-search-roi/
- https://www.conductor.com/academy/aeo-geo-benchmarks-report/
- https://developers.google.com/search/blog