Short answer. The hard part of a GEO content strategy is not how to write the page — it is deciding which pages are worth writing at all. The selection rule that holds up is: publish where the question is commercially real, where the incumbent answer is weak or generic, and where you can say something non-substitutable — first-party data, a defined method, a stated formula, an honest limit. Everything else is a page a model can already assemble from three other sources, which is exactly why it will keep using those three. Coverage is not the goal; being the only reasonable place to get a particular fact is.
Most GEO plans fail at prioritisation rather than execution. A team identifies 200 questions, writes the 40 easiest, and reports no movement — because the 40 easiest were the 40 already answered adequately elsewhere. This guide is the selection layer: how to build a question inventory, how to score it, what to publish against each score, and what to deliberately not write.
What should a GEO strategy actually optimise for?
"Get cited by ChatGPT" is not an objective, because it collapses six separable outcomes into one. A page can be indexed and never retrieved; retrieved and never selected; selected and never visibly cited; cited and send no traffic; or influence an answer without appearing at all.
| Outcome | What it means | What moves it |
|---|---|---|
| Discoverability | The page is in the index at all | Crawlability, rendering, internal links |
| Retrieval | The page enters the candidate pool for a question | Topical match to the sub-questions actually asked |
| Selection | The passage enters the model's context | Self-containment, directness, evidence |
| Citation | The source is visibly attributed | Entity resolution, unambiguous ownership of the claim |
| Mention | The brand is named without a link | Third-party corroboration as much as owned content |
| Business result | Pipeline, not visibility | Whether the question was commercially real |
Separating these is what makes a diagnosis possible. A programme reporting one blended score cannot tell the difference between "we are not in the index" and "we are in the index and boring", and those need opposite responses.
Where do the questions come from?
Not from a keyword tool, and not from an English question list translated into other markets. Three sources, in descending order of value:
- Sales and support transcripts. The questions a buyer asks a human are the questions they ask an assistant, phrased almost identically. This is the highest-yield input and the one most teams already own and never read.
- The assistants themselves. Ask category, comparison and problem questions across engines and record what comes back — which questions produce confident answers, which produce hedged ones, and who is named.
- In-market collection per language. Questions differ between markets in substance, not only in wording. A translated question list embeds the source market's assumptions about what buyers care about, and then reports honestly that the translated pages did not perform.
Record each question with the market, the language, the buying stage, and who currently answers it well. That last field is the one that drives the decision.
How do you score a question?
Four criteria. Score each 0–3; publish against the total, not against enthusiasm.
| Criterion | 0 | 3 |
|---|---|---|
| Commercial reality | Nobody who asks this buys anything | The question sits directly before a purchase decision |
| Incumbent weakness | Answered well by an authoritative source | Answers are generic, contradictory, or visibly hedged |
| Non-substitutability | We would be restating public knowledge | We hold data, a method, or an outcome nobody else can state |
| Maintainability | The answer changes monthly and nobody owns it | Stable, or a named owner exists to update it |
10–12: write it properly. Full treatment, first-party evidence, maintained. 7–9: write it if capacity allows, after the tier above is complete. 4–6: fold it into an existing page as a section rather than giving it a URL. 0–3: do not write it. This is the tier that consumes most GEO budgets.
The third row does most of the discrimination. If the honest answer to "what can we say here that another source cannot" is nothing, the page will be correct, competent and unused.
What makes a page non-substitutable?
A model can already produce a fluent overview of almost any topic. What it cannot produce is a specific, attributable fact that exists in exactly one place. Five things qualify:
- First-party measurement. Numbers you produced, with the method stated. A figure with a described method is quotable; a figure without one is a claim.
- A defined formula. Passages that state how something is calculated are unusually citable and unusually rare — most vendor sites in most categories carry no formulas at all in body copy.
- A stated threshold. "Pass at eight per thousand words" is liftable. "High quality" is not.
- A named limitation. Pages that explain where a method fails are more credible, and more useful to a system assembling a balanced answer, than pages claiming universal success.
- An outcome with its conditions attached. Not "we improved results" but what changed, over what period, measured how, and what did not move.
Google's guidance on generative AI features points the same way: unique, useful, people-first content rather than near-duplicate pages generated for every query variant. The strategic version of that guidance is this scoring table.
What about the questions you cannot win on your own site?
A significant share of the answers about any category are assembled from third-party platforms rather than vendor websites. For those, publishing harder on your own domain addresses the smaller half of the problem.
The realistic split:
- Owned content wins definitional, methodological and "how do I do X" questions, where the best available answer can genuinely be yours.
- Third-party corroboration wins "who should I use", "what are the alternatives to X" and reputation questions, which are assembled from review platforms, editorial listicles, community threads and reference sites.
- Neither wins quickly on model memory. Being described accurately by a model answering with no browsing changes at training cadence, through what exists about you elsewhere, over months.
Earned coverage has to be genuine. Planted reviews, fabricated listicles and citation farms produce short-term mentions, conflict with platform policy, and are the kind of signal that gets discounted rather than rewarded once detected.
A publishing cadence that survives contact with reality
- Fewer, maintained. A library of 25 pages that are updated beats 120 that are not, because staleness is visible and dated claims are checkable.
- One page per question, not per phrasing. Consolidate wording variants into one substantive answer. Splitting them produces near-duplicates that compete with each other and trip scaled-content policy.
- Name an owner per page. Unowned pages decay into wrong pages, and a wrong page is worse than an absent one once an assistant repeats it back to a prospect.
- Re-score annually. Incumbent weakness is the criterion that changes fastest. A question that was wide open last year may now be answered well by someone with more authority than you.
How Lifewood approaches this
Lifewood treats question selection as the first deliverable of a GEO programme, before any writing. The inventory is built from client sales and support material and from live assistant runs rather than from keyword exports, scored on the four criteria above, and the "do not write" tier is delivered explicitly — because the pages a client is talked out of are usually where the budget was going to go.
Content is then produced in-market rather than translated. With 50+ languages, 40+ delivery centres across 30+ countries and 56,788 registered contributors, the question inventory for each market is collected by people who sell into it, which is the only way the substance of the question survives the crossing.
The limit worth stating: none of this makes a page citable if the entity behind it cannot be resolved or the page cannot be fetched in full. Those are fixed first. See GEO services and AEO services.
Sources and further reading
- Google Search Central, "Optimizing your website for generative AI features" — the people-first, non-scaled-content position underpinning the selection rule.
- Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024 — why evidence rather than coverage is the lever.
- Companion guides: Where AI Answer Engine Citations Actually Go (the owned-versus-third-party split) and What to Look For in AEO and GEO Services.

