LIFEWOOD
Ready100
AI search

10 Questions to Ask Before Hiring AEO and GEO Help

Short answer. Ten questions separate an AI search visibility provider that will move something from one that will bill you for a dashboard. Ask about the baseline, the memory/retrieval…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. Ten questions separate an AI search visibility provider that will move something from one that will bill you for a dashboard. Ask about the baseline, the memory/retrieval split, raw run files, who writes the content, entity work, prompt-set authorship per language, what they expect not to move, their worst failure, data and content ownership, and what changed in the last year that invalidated their own advice. Below each question is what a strong answer sounds like, what a weak one sounds like, and what should end the meeting.

Every provider in this category can produce a competent-looking proposal. The differences that predict results are behavioural — how they measure, what they publish, what they admit — and they surface in conversation rather than in documents.

Use this as an answer key. Take it into the meeting, ask the ten in order, and score as you go. Providers who are strong on questions 1, 2 and 4 are almost always strong on the rest; providers who deflect on those three rarely recover.


1. What is your baseline procedure, and what happens if we skip it?

Why it matters. Without a measurement taken before any work begins, nothing afterwards is attributable — to them or to anyone else. A provider who skips the baseline has voluntarily given up the only clean evidence of their own value, which is a strange decision unless the evidence was never the point.

Strong answer. A fixed prompt set of roughly 20–40 questions per market, run before execution across both answer surfaces, several runs per prompt, raw output retained, delivered to you as a dated artefact.

Weak answer. "We'll pull your current visibility score from the platform in week one."

Disqualifying. "We don't really need a baseline — you'll see the difference."

2. Do you report model memory and retrieval separately? Show me a client example.

Why it matters. A model answering from its training weights moves on model-release timescales; the same model with web search enabled responds within weeks. Blended into one figure, a genuine retrieval win stays invisible for months — and that is exactly when programmes get cancelled.

Strong answer. Two lines on every chart, explained without prompting, with different expectations attached to each.

Weak answer. "Our score covers all AI surfaces."

Disqualifying. Not knowing the distinction exists.

3. Show me a raw run file from a live client period.

Why it matters. A dashboard is a computation. The run file is the evidence — prompts, timestamps, model and mode, full returned text. Reading the actual answers is also where diagnosis comes from; a number tells you something moved, the text tells you why.

Strong answer. Produced within a day, redacted for client identity, with an explanation of the schema.

Weak answer. A screenshot of a dashboard.

Disqualifying. "That's proprietary." The methodology may be; the raw output of your own prompts is not.

4. Who writes the content, and who publishes it?

Why it matters. This is the single largest hidden cost in the category. A provider who audits and advises has moved the expensive work — writing, publishing, maintaining — onto a team that is already at capacity. The backlog does not get done, and a year later nothing has moved.

Strong answer. Named writers and reviewers, a publishing cadence, and a defined path to live including who has CMS access.

Weak answer. "We'll provide detailed briefs for your team."

Disqualifying. Discovering at week six that "delivery" meant a spreadsheet of recommendations.

If any part of the answer is "you", price that work and add it to their fee before comparing proposals.

5. What would you fix in our entity signals before writing a word of content?

Why it matters. If a model cannot resolve your brand as one corroborated entity, content volume will not fix it. The diagnostic pattern is common: models answer "what does [brand] do?" correctly but never return you for "who provides [category]?" — entity known, category association missing.

Strong answer. A page of specifics from looking at your site: naming inconsistencies between schema and copy, missing or unresolvable third-party references, expertise categories absent at the entity level, regions stated as "worldwide", contradictory structured data.

Weak answer. "We'd start with a full audit." (Everyone starts with an audit. The question is what they can already see.)

Disqualifying. Going straight to a content calendar. The entity layer is the cheapest available win and skipping it is a tell.

6. Who writes the prompt sets for our non-English markets?

Why it matters. A translated prompt set measures how a market would ask if it thought in English. Buyers in different markets phrase questions differently, compare against different competitor sets, and are convinced by different evidence.

Strong answer. Named in-market native speakers, with counts per language, and the distinction drawn between people who can review and people who can write.

Weak answer. "Our platform supports 40+ languages."

Disqualifying. Prompt sets that turn out to be machine-translated from English after you ask a second time.

7. Which of our questions do you expect not to move, and why?

Why it matters. Some categories are held by decades of corpus mass — incumbents with long histories, heavy third-party coverage and encyclopaedic presence. No quarter of work displaces that. A provider who expects everything to move is either inexperienced or managing your expectations to the point of signature.

Strong answer. A specific list, with reasons, and a proposal to contest the enterable categories first while building slowly toward the hard ones.

Weak answer. "With the right strategy everything is achievable."

Disqualifying. A guarantee of placement in any AI answer. Nobody controls the output of a model they do not operate.

8. What is the worst outcome you've had, and what changed afterwards?

Why it matters. Anyone operating at real volume has had a programme that did not work. The answer reveals whether they measure honestly, and whether the organisation learns.

Strong answer. A specific case, the diagnosis, and the change to their method that followed.

Weak answer. A reframed success story.

Disqualifying. "We haven't had one." Either they are new, or they are not measuring.

9. Who owns the content, the prompt sets and the measurement data?

Why it matters. These are the assets. A programme is a compounding investment only if you keep what it produces.

Strong answer. You own all of it; export in a non-proprietary format at any time, and automatically at termination.

Weak answer. "Everything lives in our platform, which you have access to during the engagement."

Disqualifying. Content licensed rather than transferred, or measurement history that disappears at contract end — which also destroys your ability to evaluate the next provider.

10. What changed in the last year that invalidated something you previously recommended?

Why it matters. This field changes underneath everyone. The answer separates practitioners from people repeating a playbook they read.

Strong answer. A concrete example — a tactic dropped, a measurement corrected, an assumption tested and abandoned.

Weak answer. Generalities about the pace of AI.

Disqualifying. Nothing has changed. In this category, that is not stability; it is not paying attention.


Scoring the meeting

Score each answer 0–2: 2 strong, 1 weak, 0 disqualifying or evasive.

Meeting score = Σ (answer scores)   — maximum 20
Score Read
16–20 Strong candidate; proceed to a paid pilot
11–15 Capable but with a specific gap — identify which, and price it
6–10 Advisory dressed as delivery; expect to do the work yourself
0–5 Reporting tool with a retainer

Any single zero on questions 1, 2, 3 or 4 should override the total. Those four are not fixable after signature by paying more.


Red flags that need no question at all

  • A guaranteed outcome. Nobody controls a model's generated answer at query time.
  • A proprietary score with no formula. If you cannot reproduce it from raw data, it is a marketing device.
  • A volume-first plan. "Sixty articles a quarter" addresses surface area, not citability — and the published evidence points the other way. In the ACM KDD 2024 benchmark across 10,000 queries, statistics raised citation visibility by up to 40% and authoritative quotations by roughly 30%, while keyword stuffing scored −10%.
  • No mention of crawlability or rendering. If your key pages only exist after JavaScript runs, most AI crawlers receive an empty page and no content plan will help.
  • Attribution charts with no discussion of confounders. Engines change underneath every measurement.
  • A single global visibility number for a business selling in several markets.

What to do after the meeting

Run a paid pilot before committing to a retainer. A defensible pilot is: baseline on both surfaces, entity fixes, three to five pages rewritten as answer-ready, and a second measurement — roughly a quarter. Judge it on the retrieval surface, since that is the only one capable of moving in the window, and on whether the raw data supports the story in the report.


How Lifewood approaches this

Lifewood answers all ten of these in its own scoping conversations, and publishes the underlying method rather than a score: fixed prompt sets, a pre-work baseline, memory and retrieval reported separately, raw run files retained, and content written and published rather than recommended.

For multi-market brands the differentiator is who writes: 50+ languages, 40+ delivery centres across 30+ countries and 56,788 contributors mean prompt sets and content authored in-market rather than translated. Lifewood also runs this programme on its own site, which is where the failure modes described above were observed rather than imagined.

See AEO services, GEO services, AEO and GEO providers and the glossary.


Sources and further reading

  • Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024 — across 10,000 queries: statistics up to +40% citation visibility, authoritative quotations roughly +30%, fluency +15–30%, keyword stuffing −10%.
  • Companion guides: 7 Things to Look for in AEO and GEO Services and 8 Signs You Need Managed AEO Services in 2026.
  • Lifewood delivery figures (50+ languages, 40+ centres, 30+ countries, 56,788 contributors) are published on lifewood.com.

Frequently asked questions

Providers fall into four groups: specialist AEO/GEO agencies, SEO agencies with an AI practice, digital PR firms working the entity and corroboration layer, and managed AI-data and content providers such as Lifewood that combine in-house measurement with multilingual execution. The right group depends on what you are missing — if your site cannot be crawled properly, no amount of content or PR will produce a recommendation.

Services that measure and improve how often a brand appears, is cited, or is recommended inside AI-generated answers. The work spans three layers: entity resolution, answer-ready content, and a measurement instrument that runs a fixed prompt set against each engine. Providers that supply only the third are reporting tools, not services.

No. Answers are generated at query time by models the provider does not operate, and vary between runs. A competent provider raises the probability and measures the change against a baseline. Treat a guarantee as a reason to end the evaluation.

The honest comparison is not the retainer but the total: their fee plus the internal cost of any work they hand back. An advisory engagement with a low retainer that requires your team to write and publish everything is frequently the more expensive option, because the work competes with a backlog and often does not happen at all.

About a quarter. That is long enough for retrieval-surface movement on published content and short enough to stop cheaply. Do not judge memory-surface results in that window — they follow model training cycles and will show nothing regardless of how good the work is.

Treat it as disqualifying. Their methodology may reasonably be proprietary; the raw output of prompts run about your brand is not, and a provider unwilling to show it is asking you to trust a number you cannot check.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team