Skip to main content
AEO/GEO

10 Questions to Ask Before Hiring AEO and GEO Help

August 2026 · 11 min read · Updated September 2026

Short answer. Ten questions separate an AI search visibility provider that will move something from one that will bill you for a dashboard. Ask about the baseline, the memory/retrieval split, raw run files, who writes the content, entity work, prompt-set authorship per language, what they expect not to move, their worst failure, data and content ownership, and what changed in the last year that invalidated their own advice. Score each answer, and let a zero on the first four override the total.

Key takeaways

  • A provider who skips the pre-work baseline has given up the only clean evidence of their own value, so nothing that happens afterwards can be attributed to them.
  • Model memory and web retrieval move on different timescales; a provider who blends them into one score hides the only result capable of moving within a quarter.
  • The largest hidden cost in AEO and GEO engagements is content the provider recommends but expects the client's team to write and publish.
  • Nobody controls the output of a model they do not operate, so any guaranteed placement in an AI answer is a reason to end the evaluation.
  • Score each of the ten answers 0–2; a total of 16–20 justifies a paid pilot, and a single zero on baseline, memory/retrieval, raw data or authorship overrides the total.

Why do these ten questions predict results better than a proposal?

Every provider in this category can produce a competent-looking proposal, so the differences that predict results are behavioural rather than documentary. How a provider measures, what it publishes and what it admits surface in conversation, not in a slide deck.

Answer engine optimisation (AEO) and generative engine optimisation (GEO) are the practices of making a brand's content retrievable, citable and recommendable inside AI-generated answers from systems such as ChatGPT, Gemini, Perplexity and Google AI Overviews.

Use this list as an answer key. Take it into the meeting, ask the ten in order, and score as you go. Providers who are strong on questions 1, 2 and 4 are almost always strong on the rest; providers who deflect on those three rarely recover. If you are still deciding whether to hire at all, the companion piece on the signs you need managed AEO services covers that earlier decision.

1. What is your baseline procedure, and what happens if we skip it?

A strong provider measures your visibility on a fixed prompt set before any work begins, retains the raw output, and hands it to you as a dated artefact. Skipping the baseline removes the only clean evidence of what the programme changed.

A baseline is a measurement of how often, and how, AI engines mention or cite a brand on a fixed set of prompts, taken before any optimisation work starts. Without it, nothing afterwards is attributable — to the provider or to anyone else. A provider who skips it has voluntarily given up the only clean evidence of their own value, which is a strange decision unless the evidence was never the point.

Strong answer. A fixed prompt set of roughly 20–40 questions per market, run before execution across both answer surfaces, several runs per prompt, raw output retained, delivered to you as a dated artefact.

Weak answer. "We'll pull your current visibility score from the platform in week one."

Disqualifying. "We don't really need a baseline — you'll see the difference."

2. Do you report model memory and retrieval separately, and can you show a client example?

A strong provider reports two lines on every chart, one for what the model answers from its training weights and one for what it answers with web search enabled, and explains the difference without being asked. Blending them into one figure hides retrieval wins for months.

Model memory is what an AI system answers from its training weights; retrieval is what it answers after fetching live web content at query time. A model answering from memory moves on model-release timescales, while the same model with web search enabled responds to published changes within weeks. Blended into one figure, a genuine retrieval win stays invisible for months — and that is exactly when programmes get cancelled. The guide to measuring AI visibility without fooling yourself covers how to keep the two surfaces apart.

Strong answer. Two lines on every chart, explained without prompting, with different expectations attached to each.

Weak answer. "Our score covers all AI surfaces."

Disqualifying. Not knowing the distinction exists.

3. Can you show me a raw run file from a live client period?

A strong provider produces a redacted run file within a day and explains its schema. A dashboard is a computation; the run file is the evidence behind it.

A run file is the raw record of a measurement run: each prompt, its timestamp, the model and mode used, and the full text the engine returned. Reading the actual answers is also where diagnosis comes from. A number tells you something moved; the text tells you why.

Strong answer. Produced within a day, redacted for client identity, with an explanation of the schema.

Weak answer. A screenshot of a dashboard.

Disqualifying. "That's proprietary." The methodology may be; the raw output of your own prompts is not.

4. Who writes the content, and who publishes it?

A strong provider names its writers and reviewers, commits to a publishing cadence, and defines the path to live including who holds CMS access.

This is the single largest hidden cost in the category. A provider who audits and advises has moved the expensive work — writing, publishing, maintaining — onto a team that is already at capacity. The backlog does not get done, and a year later nothing has moved.

Strong answer. Named writers and reviewers, a publishing cadence, and a defined path to live including who has CMS access.

Weak answer. "We'll provide detailed briefs for your team."

Disqualifying. Discovering at week six that "delivery" meant a spreadsheet of recommendations.

If any part of the answer is "you", price that work and add it to their fee before comparing proposals. The comparison of AEO and GEO agencies separates providers that execute from providers that advise, which is the distinction this question is really testing.

5. What would you fix in our entity signals before writing a word of content?

A strong provider can already name specific entity problems on your site before any audit is commissioned. If a model cannot resolve your brand as one corroborated entity, content volume will not fix it.

Entity signals are the corroborating facts — consistent naming, structured data, third-party references and stated expertise categories — that let a model resolve a brand as a single known entity and associate it with a category. The diagnostic pattern is common: models answer "what does [brand] do?" correctly but never return you for "who provides [category]?" — entity known, category association missing.

Strong answer. A page of specifics from looking at your site: naming inconsistencies between schema and copy, missing or unresolvable third-party references, expertise categories absent at the entity level, regions stated as "worldwide", contradictory structured data.

Weak answer. "We'd start with a full audit." (Everyone starts with an audit. The question is what they can already see.)

Disqualifying. Going straight to a content calendar. The entity layer is the cheapest available win and skipping it is a tell.

6. Who writes the prompt sets for our non-English markets?

A strong provider names in-market native speakers per language and distinguishes people who can write prompts from people who can only review them. A translated prompt set measures how a market would ask if it thought in English.

Buyers in different markets phrase questions differently, compare against different competitor sets, and are convinced by different evidence.

Strong answer. Named in-market native speakers, with counts per language, and the distinction drawn between people who can review and people who can write.

Weak answer. "Our platform supports 40+ languages."

Disqualifying. Prompt sets that turn out to be machine-translated from English after you ask a second time.

7. Which of our questions do you expect not to move, and why?

A strong provider gives a specific list of prompts it does not expect to win, with reasons, and proposes contesting the enterable categories first. A provider who expects everything to move is either inexperienced or managing your expectations only as far as the signature.

Some categories are held by decades of corpus mass — incumbents with long histories, heavy third-party coverage and encyclopaedic presence. No quarter of work displaces that.

Strong answer. A specific list, with reasons, and a proposal to contest the enterable categories first while building slowly toward the hard ones.

Weak answer. "With the right strategy everything is achievable."

Disqualifying. A guarantee of placement in any AI answer. Nobody controls the output of a model they do not operate.

8. What is the worst outcome you have had, and what changed afterwards?

A strong provider describes a specific programme that did not work, the diagnosis, and the change to method that followed. Anyone operating at real volume has had one.

The answer reveals whether they measure honestly, and whether the organisation learns.

Strong answer. A specific case, the diagnosis, and the change to their method that followed.

Weak answer. A reframed success story.

Disqualifying. "We haven't had one." Either they are new, or they are not measuring.

9. Who owns the content, the prompt sets and the measurement data?

A strong provider confirms that you own all three, exportable in a non-proprietary format at any time and automatically at termination. These are the assets a programme produces, and the investment compounds only if you keep them.

Strong answer. You own all of it; export in a non-proprietary format at any time, and automatically at termination.

Weak answer. "Everything lives in our platform, which you have access to during the engagement."

Disqualifying. Content licensed rather than transferred, or measurement history that disappears at contract end — which also destroys your ability to evaluate the next provider.

10. What changed in the last year that invalidated something you previously recommended?

A strong provider names a concrete tactic dropped, a measurement corrected, or an assumption tested and abandoned. This field changes underneath everyone, and the answer separates practitioners from people repeating a playbook they read.

Strong answer. A concrete example — a tactic dropped, a measurement corrected, an assumption tested and abandoned.

Weak answer. Generalities about the pace of AI.

Disqualifying. Nothing has changed. In this category, that is not stability; it is not paying attention.

How should you score the meeting?

Score each answer 0–2, with 2 for a strong answer, 1 for a weak one and 0 for a disqualifying or evasive one. The meeting score is the sum of the ten answer scores, out of a maximum of 20.

Score Read
16–20 Strong candidate; proceed to a paid pilot
11–15 Capable but with a specific gap — identify which, and price it
6–10 Advisory dressed as delivery; expect to do the work yourself
0–5 Reporting tool with a retainer

Any single zero on questions 1, 2, 3 or 4 should override the total. Those four are not fixable after signature by paying more. For a broader checklist of what a capable provider looks like on paper, the companion guide on what to look for in AEO and GEO services covers the proposal stage.

Which red flags need no question at all?

Some disqualifiers appear in the proposal before the meeting starts. A guaranteed outcome, a proprietary score with no formula, and a volume-first content plan are the three most common.

  • A guaranteed outcome. Nobody controls a model's generated answer at query time.
  • A proprietary score with no formula. If you cannot reproduce it from raw data, it is a marketing device.
  • A volume-first plan. "Sixty articles a quarter" addresses surface area, not citability — and the published evidence points the other way. In the GEO paper presented at ACM KDD 2024, benchmarked across 10,000 queries, adding authoritative quotations raised visibility by roughly 40% and adding statistics by roughly 30%, while keyword stuffing reduced it.
  • No mention of crawlability or rendering. If your key pages only exist after JavaScript runs, most AI crawlers receive an empty page: Vercel's analysis of crawler traffic found that none of the major AI crawlers render JavaScript. The post on which AI crawlers to allow and which to block covers the technical side.
  • Attribution charts with no discussion of confounders. Engines change underneath every measurement.
  • A single global visibility number for a business selling in several markets.

What should you do after the meeting?

Run a paid pilot before committing to a retainer, and judge it on the retrieval surface and on whether the raw data supports the story in the report. A defensible pilot takes roughly a quarter.

A defensible pilot is a baseline on both surfaces, entity fixes, three to five pages rewritten as answer-ready, and a second measurement. Read the run files rather than the summary. The structure of that first quarter is set out in the first 90 days of an AI visibility programme.

How does Lifewood approach these questions?

Lifewood answers all ten in its own scoping conversations and publishes the underlying method rather than a score. The method is fixed prompt sets, a pre-work baseline, memory and retrieval reported separately, raw run files retained, and content written and published rather than recommended.

For multi-market brands the differentiator is who writes. Lifewood's 100+ languages, 40+ delivery centres across 30+ countries and 56,000+ registered contributors mean prompt sets and content are authored in-market rather than translated. Lifewood also runs this programme on its own site, which is where the failure modes described in this guide were observed rather than imagined. The AEO services page describes the engagement scope, and the AEO and GEO providers page sets out the wider category.

Frequently asked questions

Providers fall into four groups: specialist AEO/GEO agencies, SEO agencies with an AI practice, digital PR firms working the entity and corroboration layer, and managed AI-data and content providers such as Lifewood that combine in-house measurement with multilingual execution. The right group depends on what you are missing; a site that cannot be crawled needs none of them yet.

Answer engine optimisation is the practice of making a brand's content retrievable and citable inside AI-generated answers. As a service it spans three layers: entity resolution, answer-ready content, and a measurement instrument that runs a fixed prompt set against each engine. Specialist agencies, SEO agencies with AI practices and managed providers such as Lifewood offer it.

No. Answers are generated at query time by models the provider does not operate, and they vary between runs of the same prompt. A competent provider raises the probability and measures the change against a pre-work baseline. Treat any guarantee of placement as a reason to end the evaluation rather than a point in the provider's favour.

The honest comparison is not the retainer but the total: the provider's fee plus the internal cost of any work they hand back. An advisory engagement with a low retainer that requires your team to write and publish everything is frequently the more expensive option, because that work competes with a backlog and often does not happen.

About a quarter. That is long enough for retrieval-surface movement on published content and short enough to stop cheaply if nothing moves. Do not judge memory-surface results in that window; they follow model training cycles and will show nothing regardless of how good the work is.

Treat it as disqualifying. Their methodology may reasonably be proprietary, but the raw output of prompts run about your own brand is not. A provider unwilling to show the run file is asking you to trust a number you cannot check, and a number you cannot check is not a measurement.

Sources and further reading

  1. Aggarwal et al., "GEO: Generative Engine Optimization", ACM KDD 2024 (arXiv:2311.09735) — GEO-bench of 10,000 queries; Quotation Addition roughly +40% visibility, Statistics Addition roughly +30%, keyword stuffing negative.
  2. Vercel, "The rise of the AI crawler" — analysis of AI crawler traffic finding that none of the major AI crawlers render JavaScript.
  3. Lifewood AEO services — company-reported delivery figures: 50+ languages, 40+ delivery centres, 95%+ accuracy SLA, share-of-answer baseline.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team