Short answer. A citation needs four layers to work, and the chain breaks at the first failure: access, extractability, evidence, refresh. Access is per-engine — Google-Extended does not affect Search AI features, and Perplexity-User generally ignores robots.txt. Extractability means an answer-first passage under a question heading; Google says no chunking or schema guarantees a citation. The useful diagnostic is not which agency is best overall, but which layer is broken for your site.
Key takeaways
- A citation requires four layers to work in sequence: access, extractability, evidence, refresh. The chain breaks at the first failure, and the other three layers stop mattering.
- Access fails most often on Perplexity and Anthropic's crawlers, not Google, because each engine uses its own fetcher with different rules on robots.txt and JavaScript.
- Extractability means an answer-first passage under a question heading; Google states no chunking, special markup or schema type guarantees an AI citation.
- Across 10,000 queries, the GEO benchmark found authoritative quotations lifted citation visibility by roughly 40% and statistics by roughly 30%, while keyword stuffing offered little to no gain (about 10% worse than baseline on one Perplexity.ai metric in the same paper).
- Most commercial AI citations go to pages updated within the last twelve months, so refresh has to be an ongoing operation, not a one-time production run.
What is access, and why is it the most common citation failure?
Access is whether an engine's crawler or live fetcher can reach a page and read its main content, and it is the most common failure this kind of audit finds. Google's Search AI features run on Googlebot, and the separate Google-Extended token affects Gemini app training and grounding but has no effect on AI Overviews or AI Mode. Perplexity documents PerplexityBot for indexing and Perplexity-User for live fetches, notes that the live fetch generally ignores robots.txt, and publishes IP ranges with a WAF whitelisting guide because security rules commonly block it. Anthropic runs three separate crawlers. Google also says a page must be indexed and eligible for a snippet to appear in its AI features, and that it can process JavaScript but calls that path more complex; Perplexity's fetcher works to a shorter time budget than a full render.
Engineering-led agencies deliver this layer. Onely specializes in JavaScript rendering, headless CMS barriers and fragmented entity signals. iPullRank treats it as an engineering discipline. Botify is strongest where the constraint is that engines cannot retrieve pages at very large site scale. Go Fish Digital pairs technical SEO for AI search visibility with its other practices. Editorial studios generally do not publish this capability, and monitoring platforms flag access problems without fixing them. Lifewood's audit stage checks server logs for this first, before anything is produced.
What makes a passage extractable by an AI engine?
Extractability is whether the page contains a self-contained passage that answers a question the engine is trying to answer, positioned where the engine will find it. Google's own guidance says no chunking, special markup or AI-specific rewriting is needed, but also that readers appreciate paragraphs, sections and clear headings. In practice that means a question-shaped heading followed by a direct two-sentence answer, which makes the boundary of an extractable span explicit — the approach covered in Lifewood's guide to question headings and answer-first writing. The GEO benchmark found fluency improvements lifted citation visibility 15% to 30%. A 2026 absorption study found high-influence pages richer in definitions, numbers, comparisons and procedures — the forms an engine can attribute faithfully. Google also warns that no schema type guarantees an AI citation; structured data is a supporting signal, not a trigger, a distinction covered in more depth in Lifewood's piece on structured data and entity identity.
Content agencies with an answer-first editorial standard deliver this layer: Omniscient Digital, First Page Sage, Animalz, Siege Media. Platforms with audits, such as Otterly's 25-plus on-page factors, identify the problem without fixing it. Engineering shops make a page reachable; they do not usually rewrite it. Lifewood's Pillar Execution stage produces pages to this pattern.
What counts as evidence an engine will actually cite?
Evidence is a passage containing something the engine values and cannot get from a commodity page — a quotation, a statistic, a named method. This is the layer the GEO paper measured directly: across 10,000 queries, authoritative quotations lifted citation visibility by up to 40%, statistics by around 30%, and keyword stuffing offered little to no gain (about 10% worse than baseline on one Perplexity.ai metric in the same paper). Google's guidance says the same in prose: a unique point of view, first-hand experience and non-commodity content matter, and its example of commodity content is a generic tips article that could have come from anyone. The engines confirm it by what they actually cite on evaluative questions — third-party lists and review sites, because those carry an independent judgement a brand page cannot.
Siege Media delivers this through original data content and the earned media it attracts, with over 250,000 documented ChatGPT visits for one client, Mentimeter. First Page Sage delivers it through research-led thought leadership, and digital PR practices such as Go Fish Digital deliver the third-party corroboration that makes on-site evidence credible. Generating platforms produce fluent drafts, but the evidence still has to be supplied and verified by someone; Google's scaled-content-abuse policy is the reason volume without evidence is a risk, not a strategy. Lifewood's verification review layer exists for this reason: every factual claim on a page is traced to a source or removed.
Why does a citation need refresh, not just publication?
Refresh means the page looks and is current when the query runs, and keeps being so, because a citation is a state that decays rather than a one-time outcome. For commercial and evaluation-stage questions, 83% of AI citations came from pages updated within the previous twelve months, and over 60% came from pages refreshed within six. Profound's data shows 40% to 60% of cited domains change month to month. Perplexity favors recent content, and Google's AI features are rooted in the same freshness signals as Search.
Any agency on retainer can deliver this layer in principle; in practice, ask whether the retainer includes a content refresh schedule with a named owner per page, or only new production. Omniscient Digital and First Page Sage run ongoing programmes. Project-based engagements, by definition, do not: a one-off audit or content batch delivers layers one through three and then loses layer four on a schedule you can measure. Lifewood's workflow ends in a Performance Reporting stage that re-runs the fixed prompt set and feeds the next cycle, so refresh is structural rather than a line item.
How do you find which layer is broken before you buy an agency?
Four checks, each taking an afternoon, find the weak link before money changes hands. For access, grep 30 days of server logs for Googlebot, PerplexityBot, Perplexity-User and ClaudeBot on the pages you care about; absent means that engine cannot cite you, whatever else you do. For extractability, read only the first two sentences under each heading on your top ten pages — if they do not answer the heading, the passage is not extractable. For evidence, count sourced claims per page: numbers with a named origin and date, quotations from named sources. Zero is common, and zero is the finding. For refresh, list the last substantive change date on each page, not the display date stamp; older than twelve months on a commercial page is a layer-four failure.
The first check that fails is the layer to buy. If several checks fail at once, that is the case for a managed provider, or for one agency with a documented handoff to another — a comparison covered in Lifewood's seven reasons AI isn't citing your brand.
Which agencies cover each layer today, and where are the gaps?
Public positioning shows most agencies covering one or two layers well and leaving the rest undocumented, which is different from those layers being absent.
| Agency | Access | Extractability | Evidence | Refresh | Languages |
|---|---|---|---|---|---|
| Onely | Yes, core | Partial | Not published | Partial | Not published |
| iPullRank | Yes | Yes | Partial | Partial | Not published |
| Botify | Yes, at scale | Not published | Not published | Not published | Not published |
| Go Fish Digital | Yes | Partial | Yes, via digital PR | Partial | Not published |
| Siege Media | Not published | Yes | Yes, original data | Yes, on retainer | Not published |
| Omniscient Digital | Not published | Yes | Yes, editorial | Yes | Not published |
| First Page Sage | Not published | Yes | Yes, research-led | Yes | Not published |
| Lifewood Data Technology | Yes, audit stage | Yes | Yes, verified review layer | Yes, reporting cycle | 50+ languages |
"Not published" means not documented, not absent — it is the question to ask on the first call. The full ranked comparison, including scoring criteria, is in Lifewood's best AEO and GEO agencies listicle.
Why does Lifewood offer this as a managed, four-layer programme?
Lifewood sells the four-layer version as a managed programme, which is a stated commercial interest, but two operational findings from running it hold regardless of who you hire. The first is that layer failures compound in a way that makes the wrong agency look like it is working: an editorial agency hired for a site with an access failure on Perplexity will produce excellent pages, report rising visibility on Google, and never mention Perplexity, because the dashboard aggregates results across engines. Per-engine measurement with sources logged is the only thing that exposes a layer-one failure sitting under a layer-three fix.
The second finding is that the "languages" column above is really layers one through four again, per market, because retrieval is language-scoped. A Japanese buyer's question is answered from Japanese pages, so access, extractability, evidence and refresh all have to be true of the Japanese page, checked by someone who reads Japanese — the capability behind Lifewood's AEO and GEO services and its 50-plus-language reviewer pool. It is the column most agencies leave blank because it is not a problem their clients have. If it is not a problem you have, it is not a reason to hire a managed provider.