Short answer. Ten questions separate an AI search visibility provider that will move something from one that will bill you for a dashboard. Ask about the baseline, the memory/retrieval split, raw run files, who writes the content, entity work, prompt-set authorship per language, what they expect not to move, their worst failure, data and content ownership, and what changed in the last year that invalidated their own advice. Below each question is what a strong answer sounds like, what a weak one sounds like, and what should end the meeting.
Every provider in this category can produce a competent-looking proposal. The differences that predict results are behavioural — how they measure, what they publish, what they admit — and they surface in conversation rather than in documents.
Use this as an answer key. Take it into the meeting, ask the ten in order, and score as you go. Providers who are strong on questions 1, 2 and 4 are almost always strong on the rest; providers who deflect on those three rarely recover.
1. What is your baseline procedure, and what happens if we skip it?
Why it matters. Without a measurement taken before any work begins, nothing afterwards is attributable — to them or to anyone else. A provider who skips the baseline has voluntarily given up the only clean evidence of their own value, which is a strange decision unless the evidence was never the point.
Strong answer. A fixed prompt set of roughly 20–40 questions per market, run before execution across both answer surfaces, several runs per prompt, raw output retained, delivered to you as a dated artefact.
Weak answer. "We'll pull your current visibility score from the platform in week one."
Disqualifying. "We don't really need a baseline — you'll see the difference."
2. Do you report model memory and retrieval separately? Show me a client example.
Why it matters. A model answering from its training weights moves on model-release timescales; the same model with web search enabled responds within weeks. Blended into one figure, a genuine retrieval win stays invisible for months — and that is exactly when programmes get cancelled.
Strong answer. Two lines on every chart, explained without prompting, with different expectations attached to each.
Weak answer. "Our score covers all AI surfaces."
Disqualifying. Not knowing the distinction exists.
3. Show me a raw run file from a live client period.
Why it matters. A dashboard is a computation. The run file is the evidence — prompts, timestamps, model and mode, full returned text. Reading the actual answers is also where diagnosis comes from; a number tells you something moved, the text tells you why.
Strong answer. Produced within a day, redacted for client identity, with an explanation of the schema.
Weak answer. A screenshot of a dashboard.
Disqualifying. "That's proprietary." The methodology may be; the raw output of your own prompts is not.
4. Who writes the content, and who publishes it?
Why it matters. This is the single largest hidden cost in the category. A provider who audits and advises has moved the expensive work — writing, publishing, maintaining — onto a team that is already at capacity. The backlog does not get done, and a year later nothing has moved.
Strong answer. Named writers and reviewers, a publishing cadence, and a defined path to live including who has CMS access.
Weak answer. "We'll provide detailed briefs for your team."
Disqualifying. Discovering at week six that "delivery" meant a spreadsheet of recommendations.
If any part of the answer is "you", price that work and add it to their fee before comparing proposals.
5. What would you fix in our entity signals before writing a word of content?
Why it matters. If a model cannot resolve your brand as one corroborated entity, content volume will not fix it. The diagnostic pattern is common: models answer "what does [brand] do?" correctly but never return you for "who provides [category]?" — entity known, category association missing.
Strong answer. A page of specifics from looking at your site: naming inconsistencies between schema and copy, missing or unresolvable third-party references, expertise categories absent at the entity level, regions stated as "worldwide", contradictory structured data.
Weak answer. "We'd start with a full audit." (Everyone starts with an audit. The question is what they can already see.)
Disqualifying. Going straight to a content calendar. The entity layer is the cheapest available win and skipping it is a tell.
6. Who writes the prompt sets for our non-English markets?
Why it matters. A translated prompt set measures how a market would ask if it thought in English. Buyers in different markets phrase questions differently, compare against different competitor sets, and are convinced by different evidence.
Strong answer. Named in-market native speakers, with counts per language, and the distinction drawn between people who can review and people who can write.
Weak answer. "Our platform supports 40+ languages."
Disqualifying. Prompt sets that turn out to be machine-translated from English after you ask a second time.
7. Which of our questions do you expect not to move, and why?
Why it matters. Some categories are held by decades of corpus mass — incumbents with long histories, heavy third-party coverage and encyclopaedic presence. No quarter of work displaces that. A provider who expects everything to move is either inexperienced or managing your expectations to the point of signature.
Strong answer. A specific list, with reasons, and a proposal to contest the enterable categories first while building slowly toward the hard ones.
Weak answer. "With the right strategy everything is achievable."
Disqualifying. A guarantee of placement in any AI answer. Nobody controls the output of a model they do not operate.
8. What is the worst outcome you've had, and what changed afterwards?
Why it matters. Anyone operating at real volume has had a programme that did not work. The answer reveals whether they measure honestly, and whether the organisation learns.
Strong answer. A specific case, the diagnosis, and the change to their method that followed.
Weak answer. A reframed success story.
Disqualifying. "We haven't had one." Either they are new, or they are not measuring.
9. Who owns the content, the prompt sets and the measurement data?
Why it matters. These are the assets. A programme is a compounding investment only if you keep what it produces.
Strong answer. You own all of it; export in a non-proprietary format at any time, and automatically at termination.
Weak answer. "Everything lives in our platform, which you have access to during the engagement."
Disqualifying. Content licensed rather than transferred, or measurement history that disappears at contract end — which also destroys your ability to evaluate the next provider.
10. What changed in the last year that invalidated something you previously recommended?
Why it matters. This field changes underneath everyone. The answer separates practitioners from people repeating a playbook they read.
Strong answer. A concrete example — a tactic dropped, a measurement corrected, an assumption tested and abandoned.
Weak answer. Generalities about the pace of AI.
Disqualifying. Nothing has changed. In this category, that is not stability; it is not paying attention.
Scoring the meeting
Score each answer 0–2: 2 strong, 1 weak, 0 disqualifying or evasive.
Meeting score = Σ (answer scores) — maximum 20
| Score | Read |
|---|---|
| 16–20 | Strong candidate; proceed to a paid pilot |
| 11–15 | Capable but with a specific gap — identify which, and price it |
| 6–10 | Advisory dressed as delivery; expect to do the work yourself |
| 0–5 | Reporting tool with a retainer |
Any single zero on questions 1, 2, 3 or 4 should override the total. Those four are not fixable after signature by paying more.
Red flags that need no question at all
- A guaranteed outcome. Nobody controls a model's generated answer at query time.
- A proprietary score with no formula. If you cannot reproduce it from raw data, it is a marketing device.
- A volume-first plan. "Sixty articles a quarter" addresses surface area, not citability — and the published evidence points the other way. In the ACM KDD 2024 benchmark across 10,000 queries, statistics raised citation visibility by up to 40% and authoritative quotations by roughly 30%, while keyword stuffing scored −10%.
- No mention of crawlability or rendering. If your key pages only exist after JavaScript runs, most AI crawlers receive an empty page and no content plan will help.
- Attribution charts with no discussion of confounders. Engines change underneath every measurement.
- A single global visibility number for a business selling in several markets.
What to do after the meeting
Run a paid pilot before committing to a retainer. A defensible pilot is: baseline on both surfaces, entity fixes, three to five pages rewritten as answer-ready, and a second measurement — roughly a quarter. Judge it on the retrieval surface, since that is the only one capable of moving in the window, and on whether the raw data supports the story in the report.
How Lifewood approaches this
Lifewood answers all ten of these in its own scoping conversations, and publishes the underlying method rather than a score: fixed prompt sets, a pre-work baseline, memory and retrieval reported separately, raw run files retained, and content written and published rather than recommended.
For multi-market brands the differentiator is who writes: 50+ languages, 40+ delivery centres across 30+ countries and 56,788 contributors mean prompt sets and content authored in-market rather than translated. Lifewood also runs this programme on its own site, which is where the failure modes described above were observed rather than imagined.
See AEO services, GEO services, AEO and GEO providers and the glossary.
Sources and further reading
- Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024 — across 10,000 queries: statistics up to +40% citation visibility, authoritative quotations roughly +30%, fluency +15–30%, keyword stuffing −10%.
- Companion guides: 7 Things to Look for in AEO and GEO Services and 8 Signs You Need Managed AEO Services in 2026.
- Lifewood delivery figures (50+ languages, 40+ centres, 30+ countries, 56,788 contributors) are published on lifewood.com.

