LLMs as time machines: what stale AI answers mean for AEO

· The Cresia team

An LLM answers from a snapshot of the past, and the answer reads the same whether that snapshot is fresh or long out of date. Nothing in the text tells the reader how far back they've been sent. For AEO and GEO teams, that means your content can fail in an answer engine in two ways at once: old facts about you keep surfacing, and current facts about you get skipped.

Key takeaways

  • An AI answer reads the same whether its facts are current or years old, so readers can't judge freshness and you can't assume they will.
  • Your content can miss in two directions: old facts about you get repeated, and new facts about you get left out.
  • Answer engines draw on remembered training data and on pages fetched at answer time. A fresh page fixes the second faster than the first.
  • Put dates, effective periods and 'formerly' wording on volatile facts, and keep one canonical page for each.
  • Measure with a fixed question set scored as correct, stale, absent or wrong, because each failure needs a different fix.
  • Don't refresh dates without changing content, and don't judge the whole picture from a single screenshot.

What the time machine framing describes

Duane Forrester, writing in Search Engine Journal, calls LLMs time machines that don't tell you how far you went. His piece says AI answers strip out the signals people once used to judge them, and that a content stack now misses in both directions. It's a good lens, and it's worth spelling out what the signals were.

On a classic results page, you saw a domain you recognised or didn't. You saw a date next to the snippet. You saw an author, a headline that hinted at the angle, and several competing results side by side. Reading those cues took a second and most people did it without noticing. An answer engine compresses all of it into one confident paragraph. The domain, the date and the disagreement between sources are gone, or tucked behind a small citation icon that most readers never open.

Here's a concrete case. A buyer asks an assistant whether a mid-size analytics vendor integrates with their CRM. The answer says yes, with a tidy explanation. The integration was retired last year. The buyer has no way to see that the explanation came from a page written before the retirement, and the vendor's team never sees the exchange at all.

That's the time machine. The trip happened, and nobody noted the destination year.

Why stale AI answers matter for AEO and GEO

Most AEO conversations start with visibility: are we mentioned, are we cited, how often. The time machine problem adds a second question that visibility counts can't answer. When you're mentioned, is what's said still true?

Think about the facts a company changes most often. Pricing and plan names. Which integrations exist. Which regions you serve. Who runs the company. What a product is called after a rebrand. A retailer that closed half its stores, a software vendor that sunset a tier, a bank that changed its eligibility rules: each has a public record that moved on while the remembered version didn't.

The damage rarely shows up as traffic. It shows up as a sales call that opens with a correction, a support ticket about a feature you don't sell, or a shortlist you never made because an answer said you lacked something you now have. None of that appears in a rank tracker.

So treat it as a brand accuracy problem first and a traffic problem second. That ordering changes who needs to be in the room. Product marketing, support and sales all hold facts that go stale, and none of them usually owns the SEO backlog.

It also runs the other way, and this is the half teams forget. If you launched something recently, the model's memory may not contain it at all. Where the product retrieves fresh pages, you may still get found. Where it doesn't, you're invisible to that question until the memory catches up, and you can't tell from the outside which case you're in.

How answer engines choose and cite sources

The mechanics differ by product and change often, so check each vendor's own documentation and help pages for the current behaviour. The broad shape is stable enough to plan around, though.

There are two layers. The first is what the model absorbed during training. It has a cutoff, and it's blended: the model doesn't store your page, it stores patterns learned from many pages, including what others wrote about you. The second layer is retrieval, where the product searches or fetches pages at answer time and feeds them to the model. Some products retrieve on most questions. Some retrieve only when the model judges it needs to. Some let the user switch it on.

This matters because the two layers age differently. A page you publish today can enter the retrieval layer as soon as it's crawled and indexed. The memory layer moves on the schedule of the next training run, which you don't control and usually can't see.

Citations add another wrinkle. A visible citation tells you a page was retrieved and shown. It doesn't always prove that a specific claim in the answer came from that page; the model may have said the same thing from memory and the citation sits beside it. When you audit, read the cited page and check whether it supports the sentence it's attached to.

What tends to get retrieved and quoted is content that's easy to lift cleanly: a heading that states the question, a first sentence that answers it, and a passage that makes sense without the paragraphs around it. A dated, specific statement beats a vague one, because it gives the model something it can say with an anchor. That's a working assumption, not a guarantee, and it's the kind of thing to test on your own questions.

What to change on your own pages

You can't edit the model's memory. You can make the retrieval layer accurate, and you can make the record that future training runs see consistent. Start here:

  1. Date the facts, not just the page. A footer that says 'last updated' helps a little. A sentence that says 'The Team plan has included single sign-on since March' helps more, because the date travels with the claim when a passage is quoted alone.
  2. Name what changed. If a product was renamed, keep a line such as 'Formerly called Studio Suite' on the current page. Old and new names then resolve to one entity instead of two half-remembered ones.
  3. Keep one canonical page per volatile fact. Pricing, integrations, availability and eligibility each get a single home. Redirect or retire older duplicates, including campaign pages and old blog posts that state the same fact differently.
  4. Write passages that stand alone. Each section should answer one question completely, with the subject named in the passage rather than referred to as 'it' or 'this plan'.
  5. Publish a plain changelog. A page listing what was added, renamed or removed, with dates, gives retrieval something authoritative to find when someone asks whether a feature still exists.
  6. Audit the pages you don't own. Review sites, partner listings, directories, conference bios and old press coverage all feed the picture. Send correction requests and update your own profiles there.
  7. Keep structured data honest. Only change a modified date when the substance changed.

The unglamorous part is number three. Most stale answers trace back to a company that itself says two things in two places.

Who owns each kind of stale fact

A fixed ownership plan beats a one-off cleanup. Here's a starting layout you can adapt.

Fact type Where the stale version tends to live Owner Review trigger
Pricing and plans Old pricing posts, partner listings, comparison pages Product marketing Any plan or price change
Integrations and features Help docs, old launch posts, marketplace profiles Product and support Any release or retirement
Product and company names Press coverage, directories, social bios Brand or comms Any rename or rebrand
Leadership and locations About pages, event bios, third-party profiles Comms and HR Any hire, exit or move
Eligibility, terms, availability Legal pages, FAQs, regional pages Legal and regional leads Any policy change

The review trigger column is the important one. Calendar-based reviews get skipped; a change to a plan that automatically creates a task for the canonical page and the third-party listings does not. That's an operations problem as much as a content one, and it's the kind of cross-team handoff covered in what marketing operations is.

How to measure accuracy, not only mentions

Counting mentions tells you whether you're in the room. It says nothing about whether you're being described correctly. Add a second measurement built around a fixed set of questions.

Pick thirty or so questions your buyers actually ask, split across the stages of a purchase: what is this category, who offers it, does vendor X do Y, how much does it cost, how does X compare with Z. Include questions about facts you've changed in the last year, since those are where staleness hides. Run the same set on a regular rhythm across the assistants your buyers use, and log the date each time.

Score every answer into one of four buckets:

  • Correct: the claim is true today.
  • Stale: the claim was true once and no longer is.
  • Absent: you or the relevant fact are missing where they should appear.
  • Wrong: the claim was never true.

Keep stale and absent separate. Stale means something old is still out there, so you look for the pages feeding it. Absent means the fact isn't reaching the answer, so you look at whether the page is retrievable, clear and consistent with what others say. Blending them into one 'accuracy score' hides which lever to pull.

Answers vary from run to run, so a single result proves little. Repeat each question a few times and look at the pattern. Record the cited source and its date when one is shown, since a stale answer with a recent cited page and a stale answer with an old cited page point to different causes.

Outside the question set, mine what you already have: notes from sales calls, support tickets that mention things you don't sell, and any referral traffic that arrives from assistants. Consistent campaign and referral tagging makes that last one readable, and it's where an analytics governance tool like OmniSpec fits. For tracking how you appear in AI search alongside your other media, see MediaPilot.

Expect this to be slow. A correction can show up in retrieval-backed answers within days and in memory-driven answers much later, if at all in the next model version. Growth teams running this as a recurring program rather than a project tend to do better, and the growth teams solution page describes how those teams are usually organised.

What not to do

The time machine framing invites some bad reactions. Skip these.

  • Don't bump dates without changing content. A fresh timestamp on an unchanged page adds nothing a reader or a retrieval system can use, and it erodes trust in your dates when the real ones matter.
  • Don't judge from one screenshot. Someone finds a wrong answer, it circulates in Slack, and the team ships a panicked rewrite. Run the question several times first. One bad answer is a lead, not a finding.
  • Don't build a page for every prompt. Dozens of thin pages that each answer a single phrasing dilute your canonical pages and create the very inconsistency that produces stale answers.
  • Don't assume fixing your site fixes the memory. It fixes retrieval. Set expectations with stakeholders that memory-driven answers lag, and that you don't control the schedule.
  • Don't hide old facts. Deleting a retired feature's page can leave only third-party descriptions behind. A short page saying what it was and when it ended is usually better than a 404.

The common thread is restraint. This is a maintenance discipline, not a campaign.

Frequently asked questions

Can I make an LLM update what it knows about my company?

Not directly. You can fix the pages that retrieval-backed answers pull from, correct third-party listings, and keep your facts consistent so later training data reflects them. Where an assistant has a feedback or correction route, check the vendor's own documentation for how it handles reports.

Is a stale answer a bigger risk than not being mentioned?

Often, yes, because a wrong claim can shape a buyer's expectations while an omission just loses a shortlist slot. Which one hurts more depends on the fact. A wrong price or a retired feature is usually worse than a missing blog post.

How often should we audit AI answers about us?

A regular rhythm that fits your release cadence works better than a fixed date. Run the full question set after any pricing, naming or product change, and a lighter pass on a set schedule. Keep the same questions each time so results are comparable.

Does adding a last-updated date help AI citations?

A visible date helps readers and gives quoted passages a time anchor, but it only works if it's honest. Dates on individual claims do more than a page-level stamp, and changing a date without changing the content is worse than leaving it alone.

Sources

  • https://www.searchenginejournal.com/llms-are-time-machines-that-dont-tell-you-how-far-you-went/590090/
  • https://searchengineland.com/topic/generative-engine-optimization
  • https://developers.google.com/search/blog

See Cresia on your own use case.

30 minutes, your team, your questions.

Request a demo