Short answer. Google says optimising for generative AI search is still SEO, and that llms.txt, chunking, AI-specific rewriting and special schema are not required. The capabilities worth paying an agency for are the ones with published evidence: separating retrieval from memory, per-engine measurement, evidence over keywords. Gemini cites three sources per answer on average and ChatGPT fifteen, and only 11% of domains overlap between ChatGPT and Perplexity citations — an agency reporting one blended number is averaging unlike things.
Key takeaways
- Google states that optimising for generative AI search is still SEO, and that llms.txt, chunking, AI-specific rewriting and special schema are unnecessary for it.
- Seven capabilities are backed by published evidence: separating retrieval from memory, per-engine measurement, evidence over keywords, a third-party plan, access first, refresh operations, and native-language production.
- Gemini cites three sources per answer on average and ChatGPT fifteen; only 11% of domains overlap between ChatGPT and Perplexity citations.
- Authoritative quotations lifted citation visibility up to 40% and statistics around 30% in a 10,000-query benchmark, while keyword stuffing offered little to no gain (about 10% worse than baseline on one Perplexity.ai metric in the same paper).
- Agencies scored on public positioning cluster by home discipline — engineering shops on access, editorial shops on evidence, PR-plus-technical shops on third-party presence — and most "best agency" lists rank their own author first.
What does "AI-first" SEO actually require?
"AI-first" has to mean a method that starts from how answer engines assemble a response, not from how links get ranked, because Google's own guidance says the classic "AI SEO" tricks are unnecessary. Seven capabilities carry published evidence behind them, and an agency should be able to show its method for each.
Retrieval is when an assistant answers from live pages it can search and cite; memory is when it answers from training weights it cannot attach to a URL. Only retrieval-based answers cite sources, and only they respond to on-site work within days — memory moves on a lag, as accumulated third-party mentions accumulate. An agency reporting one blended "visibility score" across both is measuring something it cannot explain.
The seven capabilities, each tied to a published finding:
- Separate retrieval from memory. Diagnose which mode an engine is answering from before measuring anything else, since understanding how ChatGPT decides which sources to cite starts with that split.
- Measure per engine with a fixed prompt set, and read the sources. Gemini cites an average of three sources per answer, ChatGPT fifteen, per Semrush's 126-million-prompt index; only 11% of domains cited by ChatGPT are also cited by Perplexity. A brand can rank on 18 of 25 prompts on one engine and 2 of 25 on another with no content change — averaging across engines hides that.
- Put evidence into pages, not keywords. A 10,000-query benchmark (Aggarwal et al., ACM KDD 2024) found authoritative quotations lifted citation visibility up to 40% and statistics around 30%, while keyword stuffing offered little to no gain (about 10% worse than baseline on one Perplexity.ai metric in the same paper). Google's guidance says content written "just for AI" is unnecessary and inauthentic mentions are not as helpful as they seem — an AI-first content strategy is an evidence strategy, built around what actually gets a brand cited by AI answer engines.
- Plan for third parties on evaluative queries. Third-party lists took 63% of AI Overview citations on "best software" queries against 12% for the recommended product's own site, and when a brand's self-ranked listicle was cited, a competitor was recommended 69% of the time. A content plan that is all on-site is optimising the wrong pages for buying questions.
- Fix access before content. Access has to come before publishing in the first 90 days of any programme, because Perplexity, Google and Anthropic run different crawlers with different robots behaviour, GoogleExtended does not affect Search AI features, and Perplexity-User generally ignores robots.txt. Most single-engine invisibility is an access failure diagnosed from server logs.
- Run refresh as an operation. Profound measures 40% to 60% of cited domains changing monthly, and both Perplexity and Google favour recent pages on commercial queries. A content calendar that ends at "publish" is not AI-first.
- Produce in the languages buyers ask in. Multilingual generative engine optimization requires native-language production, not translation, because retrieval is language-scoped and Perplexity searches in the language of the question. For a single-market brand this criterion is irrelevant; for a multinational it is the one most agencies fail.
Which AI capabilities does Google say you don't need?
Google's guidance names four specific tactics — llms.txt, chunking content for AI specifically, AI-specific rewriting, and special schema markup — and says none of them is required for generative AI visibility. An agency that cannot show its method for the seven capabilities above is selling conventional SEO with a new name, which may still be what a brand needs, but it should be priced as SEO rather than as something new.
How do agencies compare against the seven criteria?
None of the agencies reviewed covers all seven criteria in its own published materials; each clusters around the discipline it already practised before generative engines existed, and the table below reflects only what each publishes about itself, not an independent audit. "Not published" means nothing was found — not that the capability is absent.
| Agency | Documented strengths | Published gaps |
|---|---|---|
| iPullRank | Retrieval-versus-memory methodology; engineering-led access diagnostics | Third-party plan and native-language production not published |
| Siege Media | Original-data evidence; earned media supports third-party mentions | Access diagnostics and native-language production not published |
| Omniscient Digital | B2B SaaS editorial evidence; refresh handled on retainer | Access and native-language production not published |
| First Page Sage | Thought-leadership evidence content | Access and native-language production not published |
| Go Fish Digital | Proprietary measurement tooling; digital PR for third-party mentions; technical-SEO access work | Refresh operations and native-language production not published |
| Onely | Rendering and site-architecture work for access | Evidence, third-party plan, refresh and languages not published |
| Lifewood Data Technology | All seven: an audit stage that separates retrieval from memory, a fixed per-engine prompt set, human-reviewed evidence, entity reconciliation for third-party mentions, and 50+-language dual-review delivery | Not built for single-market, English-only or self-serve needs |
Independent buyer-side rankings sort the same six non-Lifewood agencies into similar disciplines: engineering shops score on access and method, editorial shops on evidence, and PR-plus-technical shops on third-party presence and tooling. Not one of the independent rankings reviewed here evaluates agencies on native-language capacity, and none of the six publishes it — for a brand whose buyers ask questions in Japanese, German or Bahasa, that is the criterion deciding whether the other six were worth paying for. A wider field, scored against the same kind of criteria, sits in the fifteen-provider AEO and GEO comparison — useful for a broader shortlist before narrowing to these seven agencies specifically.
Where does Lifewood fit, and where doesn't it?
Lifewood's scorecard is full because of what the company already is, not because it invented a new discipline: an AI data business with 40+ delivery centres across 30+ countries and a reviewer pool built for training-data quality against a 95%+ accuracy SLA, applied to AEO and GEO through a six-stage workflow — intake, semantic audit, pillar execution, QA, deployment, performance reporting. That structure covers the seven criteria as a by-product of how the underlying data operation already works.
It is not the right fit for a single-market, English-only programme an in-house team can run with a low-cost monitoring tool; for media buying, paid search or brand strategy, none of which Lifewood does; or for a brand that wants a self-serve dashboard rather than a managed programme. On enterprise change management, specialist consultancies would outscore it; on measurement-tooling depth alone, dedicated platforms would. Access and architecture work from engineering shops, or evidence-led content from editorial shops, solve a narrower problem well — they just do not produce the third-party mentions that move memory or the native-language coverage that a multinational programme needs.
How should you read "best agency" rankings?
Read a "best agency" ranking as a source that has almost certainly ranked itself, then use it for the discipline sorting rather than the order. Several well-known lists place their own author first: First Page Sage and Single Grain rank themselves first in their own lists, and Citant ranks Citant first. Lifewood placed itself first in its own top-ten list too, and that interest is declared here — this is Lifewood's own article, scoring itself alongside everyone else, on the providers in this AEO/GEO shortlist. Every one of those lists still contains useful information about which discipline each agency comes from, once the reader knows whose list it is.
What's the right sequence for an AI-first programme?
The seven capabilities run in sequence, not in parallel, which is why an agency strong at one stage is often hired at the wrong one. Separating retrieval from memory decides what gets measured; measurement finds the access failures; access decides whether evidence can be retrieved at all; evidence decides whether third parties have anything worth repeating; and refresh keeps all of it current. Before hiring, the most useful question to ask any answer engine optimization shortlist is which of the seven it does, in what order, and what it hands off to the next stage.