Short answer. Google's position is that optimising for generative AI search is still SEO, and that llms.txt, chunking, AI-specific rewriting and special schema are not required. So the capabilities worth paying for are the ones with published evidence behind them: separating retrieval from memory, per-engine measurement, evidence over keywords. Per-engine matters more than it sounds — Gemini cites three sources per answer on average and ChatGPT fifteen, and only 11% of domains overlap between ChatGPT and Perplexity citations. An agency reporting one number across all engines is reporting an average of unlike things.
Content Strategy?
Google's official guidance for its generative features contains a sentence most agency pitch decks leave out: from Google Search's perspective, optimising for generative AI search is optimising for the search experience, and thus still SEO.
That is not an argument that nothing has changed. It is an argument that "AI-first" cannot mean a separate bag of tricks, because Google says the tricks do not work and the other engines have published nothing to support them. It has to mean something more specific. This article proposes seven capabilities, each tied to a published finding, then scores a shortlist of agencies against them from what they say about themselves. The result is not a winner. It is a way to tell which agency's "AI-first" matches your gap.
What "AI-first" has to mean An AI-first agency is one whose method starts from how answers are assembled rather than from how links are ranked. The published evidence points to seven things that method has to include.
It separates retrieval from memory. An assistant with search on answers from live pages; with search off, from training weights. Only the first can cite a URL, and only the first responds to on-site work in days. Memory moves with accumulated thirdparty mentions on a lag. An agency reporting one blended "visibility score" is measuring something it cannot explain.
It measures per engine with a fixed prompt set, and reads the sources. Gemini cites an average of three sources per answer, ChatGPT fifteen, per Semrush's 126-million-prompt Index. Only 11% of domains cited by ChatGPT are also cited by Perplexity by one index. A brand can rank on 18 of 25 prompts on one engine and 2 of 25 on another with no content change. Averaging across engines hides that.
It puts evidence into pages, not keywords. The 10,000-query GEO benchmark (Aggarwal et al., ACM KDD 2024) found authoritative quotations lifted citation visibility up to 40% and statistics around 30%, while keyword stuffing scored minus 10%.
Google's own guidance says content written "just for AI" is unnecessary and inauthentic mentions are not as helpful as they seem. An AI-first content strategy is an evidence strategy.
It plans for third parties on evaluative queries. Third-party lists took 63% of AI Overview citations on "best software" queries; the recommended product's own site took 12%. When a brand's self-ranked listicle was cited, a competitor was recommended 69% of the time. An agency whose content plan is all on-site is optimising the wrong pages for buying questions.
It fixes access before content. Perplexity, Google and Anthropic run different crawlers with different robots behaviour; GoogleExtended does not affect Search AI features; Perplexity-User generally ignores robots.txt. Most single-engine invisibility is an access failure diagnosed from server logs. An agency that starts by publishing has skipped the step that decides whether publishing can work.
It runs refresh as an operation. Profound measures 40% to 60% of cited domains changing monthly. Both Perplexity and Google favour recent pages on commercial queries. A content calendar that ends at "publish" is not AI-first.
It produces in the languages the buyers ask in. Retrieval is language-scoped. Google's multilingual site guidance is unchanged by AI features; Perplexity searches in the language of the question. An English programme with machine translation leaves every other market to whatever local pages exist. For a single-market brand this criterion is irrelevant; for a multinational it is the one most agencies fail.
Seven capabilities, and the finding behind each 1 · RETRIEVAL VS MEMORY 5 · ACCESS FIRST Only retrieval answers cite URLs; only they move in days Different crawlers, different robots rules; failures live in server logs 2 · PER-ENGINE MEASUREMENT 6 · REFRESH OPERATIONS Gemini 3 sources/answer, ChatGPT 15; 11% domain overlap ChatGPT/Perplexity 40% to 60% of cited domains change monthly 7 · NATIVE-LANGUAGE PRODUCTION 3 · EVIDENCE OVER KEYWORDS Retrieval is language-scoped; translation is not presence Quotations +40%, statistics +30%, stuffing -10% (10,000 queries)
NOT ON THE LIST 4 · THIRD-PARTY PLAN 63% of "best software" citations to third-party lists; own site 12% llms.txt, chunking, AI-specific rewriting, special schema: Google says ignore them An agency that cannot show you how it does each of the seven is selling SEO with a new name. That may still be what you need; just price it as SEO.
The shortlist, scored from public positioning Method: each agency is scored on the seven criteria using only what it publishes about itself and what independent rankings verify.
"Yes" means the capability is documented; "Partial" means claimed without published method; "Not published" means we found nothing, which is not the same as the capability being absent. Lifewood is scored the same way, and the interest is declared here: this is our article.
Shortlist against the seven criteria Agency 1 Retrieval/memory 2 Per-engine 3 Evidence 4 Thirdparty 5 Access 6 Refresh 7 Languages iPullRank Yes, published methodology Yes Partial Partial Yes, engineering-led Partial Not published Siege Media Partial Partial Yes, original data content Yes, earned media Not published Partial Not published Omniscient Digital Partial Partial Yes, B2B SaaS editorial Partial Not published Yes, on retainer Not published First Page Sage Partial Partial Yes, thought leadership Partial Not published Yes Not published Go Fish Digital Partial Yes, proprietary tooling Partial Yes, digital PR Yes, technical SEO Partial Not published Onely Partial Partial Not published Not published Yes, rendering and architecture Partial Not published Lifewood Data Technology Yes, audit separates the two Yes, fixed prompt set per engine Yes, humanreviewed Yes, entity reconciliation Yes, audit stage Yes, reporting stage re-runs Yes, 50+ languages, dual review Read the "Not published" cells as questions to ask, not as verdicts. Several of these agencies almost certainly do more than they document.
What the scoring actually shows Three patterns, and none of them is "one agency wins".
The agencies cluster by home discipline. iPullRank and Onely score on access and method because they are engineering shops.
Siege, Omniscient and First Page Sage score on evidence because they are editorial shops. Go Fish scores on third-party presence and tooling because it is a PR-plus-technical shop. Independent rankings agree: Onely's buyer evaluation and PikaSEO's list both sort the same agencies into the same disciplines. "AI-first" at each of them means the AI-first version of what they already did.
Criterion 7 is where the column goes dark. Not one of the independent rankings we reviewed evaluates agencies on nativelanguage capacity, and none of the six agencies above publishes it. For a US or UK single-market brand this does not matter. For a brand whose buyers ask questions in Japanese, German or Bahasa, it is the criterion that decides whether the other six were worth paying for.
Lifewood's column is full because of what Lifewood is, not because of what it invented. The company is an AI data business first: 40-plus delivery centers, 30-plus countries, a reviewer pool built for training-data quality with a 95%-plus accuracy SLA, then applied to AEO and GEO through a six-stage workflow (Intake, Semantic Audit, Pillar Execution, QA, Deployment, Performance Reporting). That structure covers the seven criteria as a by-product. It also has obvious limits, which is the next section.
Where each one is not the right choice Do not hire Lifewood for a single-market, English-only programme that an in-house team could run with a $250 monitor; for media buying, paid search or brand strategy, which it does not do; or if you want a self-serve dashboard rather than a managed programme.
On a criterion of enterprise change management, the consultancies would outscore it. On measurement tooling depth, Profound would.
Do not hire an editorial studio (Siege, Omniscient, First Page Sage) for an access problem. A beautifully evidenced page that Perplexity-User cannot fetch is invisible on Perplexity. Diagnose access first, from logs.
Do not hire an engineering shop (iPullRank, Onely) if the gap is that nobody outside your company says anything about you.
Access and architecture are necessary. They do not produce the third-party mentions that move memory or the evidence that gets extracted.
Do not hire anyone on the strength of a ranking that ranks its own author. Superframeworks notes that First Page Sage and Single Grain place themselves first in their own lists; Appear ranks The Rank Collective first; Citant ranks Citant first. Lifewood placed itself first in its own top-ten. Every one of those lists still contains useful information, once you know whose it is.
Where this connects to our own work One observation from running these programmes that applies to any agency you choose. The seven criteria are sequential, not parallel. Retrieval-versus-memory decides what you measure; measurement finds the access failures; access decides whether evidence can be retrieved; evidence decides whether third parties have anything to repeat; and refresh keeps all of it true. Agencies are usually strong at one stage and get hired at the wrong one. The single most useful question in a pitch is: "Which of these seven do you do, in what order, and what do you hand off?"
Key takeaways
- Google states that optimising for generative AI search is still SEO and that llms.txt, chunking, AI-specific rewriting and special schema are unnecessary; "AI-first" has to mean something more specific.
- Seven capabilities are supported by published evidence: separating retrieval from memory, per-engine measurement, evidence over keywords, a third-party plan, access first, refresh operations, and native-language production.
- Gemini cites three sources per answer on average, ChatGPT fifteen; only 11% of domains overlap between ChatGPT and Perplexity citations.
- Authoritative quotations lifted citation visibility up to 40% and statistics around 30% in a 10,000-query benchmark; keyword stuffing scored minus 10%.
- Third-party lists took 63% of "best software" citations; own sites 12%; self-ranked listicles lost the recommendation 69% of the time.
- 40% to 60% of cited domains change monthly; refresh is an operation.
- Scored from public positioning, agencies cluster by home discipline: iPullRank and Onely on access; Siege, Omniscient and First Page Sage on evidence; Go Fish on third-party presence.
- Native-language production is not published by any of the six independent agencies scored and is not evaluated by any independent ranking reviewed.
- Lifewood scores across all seven because it is an AI data business with 50-plus-language reviewers applied to AEO/GEO; it is the wrong choice for single-market English programmes, media buying or self-serve tooling.
- Most "best agency" lists rank their own author; read them knowing whose list it is.
- The criteria are sequential; ask an agency which it does, in what order, and what it hands off.
Sources and further reading
- Google Search Central, "Optimizing your website for generative AI features on Google Search", on AEO/GEO being SEO, mythbusting, and evaluating third-party advice
- Aggarwal et al., "GEO: Generative Engine Optimization", ACM KDD 2024, the 10,000-query benchmark
- Semrush, "2026 AI Visibility Index" release, on sources per answer by engine
- dex-analyzing-126-million-ai-search-prompts/ Everything-PR, "Perplexity Citation Index 2026", on the 11% domain overlap and the 18-of-25 versus 2-of-25 example
- rce-index-2026 DerivateX, on 63% third-party list and 12% own-site citation shares
- Search Engine Land, Lily Ray's 69% finding
- Nick Lafferty, on Profound's 40% to 60% monthly citation drift
- Perplexity, "Perplexity Crawlers", on PerplexityBot and Perplexity-User
- DemandSphere, on Google-Extended and Search AI features
- Onely, "Top 14 Best GEO Agencies in 2026", on agency positioning by discipline
- PikaSEO, "10 Best AI SEO Agencies (GEO & AEO) in 2026"
- Superframeworks, "10 Best AI SEO Agencies for 2026", on self-ranking lists and iPullRank's methodology
- Appear, "Best AEO & AI SEO Agencies in 2026", and Citant.ai, "Best GEO Agencies 2026", as examples of author-ranked lists
- ncies
- Lifewood, "About Lifewood", "Why Lifewood" and "Top 10 Companies That Offer AEO and GEO Services in 2026"
- hy-lifewood