Short answer. Judge a ChatGPT visibility partner on whether they can separate the two surfaces ChatGPT answers from — model memory (training weights, which move on model-release timescales) and retrieval (live web search, which moves in weeks) — and show you results for each. Then check three things: a fixed prompt set with a pre-work baseline and repeated runs, in-house execution capacity rather than a recommendations deck, and honesty about what cannot be controlled. A partner selling a single blended "AI visibility score" and a guarantee of ChatGPT mentions has failed the first test.
ChatGPT brand visibility is now a procurement category, which means it has attracted the full range of suppliers — from disciplined AEO practices to dashboards with a markup. The category is unusually hard to evaluate because the outcome is stochastic, the mechanics are partly undocumented, and almost nobody in the buying organisation can independently verify a claim.
This guide is for the evaluation itself: what a competent partner's method looks like, what evidence to demand, how to score, and what to write into the contract. If you want the in-house playbook instead — the work itself rather than who does it — see the companion guide on improving ChatGPT brand visibility.
What are you actually buying?
Two surfaces, two timescales, two different kinds of work. Every serious conversation with a partner starts here.
| Model memory | Retrieval | |
|---|---|---|
| What the answer draws on | Training data baked into the weights | Live web search at answer time |
| Time to move | Model generations — months | Days to weeks |
| What moves it | Broad corpus presence: third-party coverage, references, mentions across the web over time | Answer-ready, crawlable, evidence-dense pages the search layer can retrieve |
| What a vendor can influence in one quarter | Very little, honestly | A great deal |
| How it is measured | Same prompts, browsing off | Same prompts, browsing on |
The practical consequence: a partner who reports one blended number cannot show you a retrieval win, because memory inertia will swamp it. Programmes get cancelled at exactly the moment they begin working, on the strength of a metric that was never able to detect the win. Require the split in the first meeting; the answer tells you most of what you need to know.
A second consequence for scoping: if your realistic horizon is a quarter, you are buying retrieval work. Memory-surface effects are a byproduct of sustained presence and third-party corroboration, and any partner promising them on a quarterly timeline is describing something outside their control.
What does a competent method look like?
Six components. Ask the partner to describe their method unprompted and check how many appear.
A fixed prompt set. Typically 20–40 questions in each in-scope language: category questions ("who provides X for enterprises"), comparison questions, and brand questions. Fixed across periods, or nothing is comparable.
A pre-work baseline. Run before any execution starts. Without it, later improvement cannot be attributed to the work — and the partner has removed the only clean evidence of their own value, which is a strange choice to make voluntarily.
Repeated runs and reported variance. Generative answers vary between runs. One run per prompt is an anecdote. Ask for runs-per-prompt and whether variance is reported alongside the mean.
Explicit metric definitions. At minimum:
Share of answer = Answers mentioning the brand ÷ Total answers for the prompt set
Cited share = Answers linking or attributing the brand ÷ Answers mentioning the brand
Mean rank = Average position where the answer returns a ranked list
Three different questions. A partner using them interchangeably is not measuring carefully.
Entity work before content work. If a model cannot resolve your brand as a single, corroborated entity, content volume will not fix it. Consistent naming across the estate, declared alternate names and transliterations, resolvable third-party references, and non-contradictory structured data come first.
An execution capability, not a recommendations deck. Ask who writes and publishes. If the answer is "your team, from our brief", price the work you are about to absorb and score them accordingly.
What evidence should you demand before signing?
Six artefacts. Each is cheap for a competent partner to produce and impossible to fake convincingly.
- A raw run file from a live client period — prompts, timestamps, model and mode, full answer text. Redacted for client identity is fine.
- A before/after pair for one prompt where they moved the outcome, with the dates of the intervening work.
- The prompt set they would use for your category, drafted before the contract. This shows whether they understand how your buyers ask.
- An entity audit of your current estate — the naming inconsistencies, missing corroboration and schema contradictions they can see from outside. A partner who cannot produce a page of specifics here has not looked.
- Two published assets they wrote, with the passage that got cited identified.
- A written statement of what they cannot control. The best partners produce this without being asked.
How should you score partners?
| Criterion | Weight | Evidence |
|---|---|---|
| Measurement method | 30% | Baseline procedure, fixed prompt set, runs per prompt, memory/retrieval split, raw file |
| Execution capacity | 25% | In-house writers and reviewers; named counts per language |
| Entity and technical foundation | 15% | Entity audit of your estate; crawl and rendering check |
| Answer-ready content quality | 15% | Published examples; evidence density; liftable passages |
| Reporting and data ownership | 10% | Raw data export; content ownership at contract end |
| Commercial terms | 5% | Notice, ramp, peak capacity |
Disqualifying, at any total score: a guarantee of ChatGPT mentions or rankings; a single blended visibility score with no surface split; no pre-work baseline; refusal to show a raw run file; a prompt set that turns out to be machine-translated for non-English markets.
Weighted score = Σ (criterion score ÷ 5 × weight)
Run the disqualification pass first. A high total score with a disqualifying condition is a well-presented version of the wrong purchase.
What belongs in the contract?
Five clauses that prevent the common disputes.
- Baseline as a deliverable. Named, dated, delivered before execution begins, with the raw file included.
- Reporting specification. Which metrics, split by surface and by market, at what cadence, with raw data export in a non-proprietary format.
- Content ownership. You own everything produced, including prompt sets and measurement data, and receive it in a usable form at termination.
- Model and method change notification. If the partner changes the model or mode used for measurement, they must say so — otherwise your trend line breaks silently and looks like a performance change.
- No guarantee clause, stated positively. A short paragraph recording that placement in generative answers is not controllable and that the engagement is measured on defined leading indicators. This protects both sides and quietly filters out the partners who will not sign it.
What are the red flags?
Guaranteed placement. Nobody controls the output of a model they do not operate.
A proprietary score with no formula. If you cannot reproduce the number from the raw data, it is a marketing device.
Volume-first proposals. "We'll publish sixty articles a quarter" addresses surface area, not citability. The published research points the other way: in the ACM KDD 2024 benchmark across 10,000 queries, statistics raised citation visibility by up to 40% and authoritative quotations by roughly 30%, while keyword stuffing scored −10%.
Language coverage counted in tool support. Ask for in-market writer and reviewer headcount per language instead.
No mention of the entity layer. A partner who goes straight to content has skipped the cheapest available win.
Attribution charts with no confounder discussion. Engines change underneath the measurement. A partner who never mentions this is either not measuring long enough to have noticed, or is choosing not to say.
Reluctance to name what will not work. Every real programme has parts that do not move. A partner who cannot name any is selling certainty rather than method.
How Lifewood approaches this
Lifewood runs ChatGPT visibility inside a single AEO and GEO programme, with the measurement instrument built and operated in-house rather than resold — which is why memory and retrieval are reported separately by default: a programme that cannot distinguish the two cannot tell a slow win from a failure.
Execution is in-house rather than briefed back to the client. 50+ languages, 40+ delivery centres across 30+ countries and 56,788 contributors mean prompt sets and published content authored by in-market native speakers, including in low-resource languages where most providers fall back to machine translation. Lifewood's AI-data heritage runs to 2004, with the current company established in 2018; engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement.
See AEO services and GEO services for scope, and the glossary for definitions of share of answer, entity canonicalization and related terms.
Sources and further reading
- Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024 — across 10,000 queries: statistics up to +40% citation visibility, authoritative quotations roughly +30%, fluency +15–30%, keyword stuffing −10%.
- Companion guide: How to Improve ChatGPT Brand Visibility in 2026 — the in-house playbook for the work described here.
- Lifewood delivery figures (50+ languages, 40+ centres, 30+ countries, 56,788 contributors) are published on lifewood.com.

