Short answer. Judge a ChatGPT visibility partner on whether they separate the two surfaces ChatGPT answers from — model memory, which moves on model-release timescales, and retrieval, which moves in weeks — and report results for each. Then check three things: a fixed prompt set with a pre-work baseline and repeated runs, in-house execution capacity rather than a recommendations deck, and honesty about what cannot be controlled. One blended "AI visibility score" plus a guarantee of mentions fails the first test.
Key takeaways
- ChatGPT answers draw on two surfaces: model memory, baked into training weights and slow to move, and retrieval, live web search at answer time, which a partner can influence within a quarter.
- A competent ChatGPT visibility method has six components: a fixed prompt set, a pre-work baseline, repeated runs with reported variance, explicit metric definitions, entity work before content work, and in-house execution.
- Six artefacts are cheap for a competent partner to produce and hard to fake: a raw run file, a before/after pair, a draft prompt set, an entity audit, two published assets with the cited passage marked, and a written statement of what they cannot control.
- Disqualify any partner, whatever their total score, who guarantees ChatGPT mentions, reports one blended visibility score, skips the baseline, or refuses to show a raw run file.
- Five contract clauses prevent the common disputes: baseline as a deliverable, a reporting specification, content ownership, model-change notification, and a positively stated no-guarantee clause.
Why is a ChatGPT visibility partner hard to evaluate?
ChatGPT brand visibility is now a procurement category, and it is unusually hard to evaluate because the outcome is stochastic, the mechanics are partly undocumented, and almost nobody in the buying organisation can independently verify a claim.
The category has attracted the full range of suppliers, from disciplined AEO practices to dashboards with a markup. A partner's claims cannot be checked against a public ranking the way a search position can, so the evaluation has to rest on method, evidence and contract terms rather than on promised outcomes.
This guide is for the evaluation itself: what a competent partner's method looks like, what evidence to demand, how to score, and what to write into the contract. If you want the in-house playbook instead — the work itself rather than who does it — the companion guide on improving ChatGPT brand visibility covers the execution side. For a shortlist of named providers, start with the top ChatGPT and AI assistant visibility companies.
What are you actually buying?
You are buying work on two surfaces with two timescales: model memory, which moves over model generations, and retrieval, which moves in days to weeks. Every serious conversation with a partner starts by separating the two.
Model memory is the knowledge baked into a model's training weights, which changes only when a new model version is trained and released. Retrieval is the live web search a model performs at answer time, which changes as soon as the search layer can find and prefer a new page.
| Dimension | Model memory | Retrieval |
|---|---|---|
| What the answer draws on | Training data baked into the weights | Live web search at answer time |
| Time to move | Model generations — months | Days to weeks |
| What moves it | Broad corpus presence: third-party coverage, references, mentions across the web over time | Answer-ready, crawlable, evidence-dense pages the search layer can retrieve |
| What a vendor can influence in one quarter | Very little, honestly | A great deal |
| How it is measured | Same prompts, browsing off | Same prompts, browsing on |
The practical consequence: a partner who reports one blended number cannot show you a retrieval win, because memory inertia will swamp it. Programmes get cancelled at exactly the moment they begin working, on the strength of a metric that was never able to detect the win. Require the split in the first meeting; the answer tells you most of what you need to know.
A second consequence for scoping: if your realistic horizon is a quarter, you are buying retrieval work. Memory-surface effects are a byproduct of sustained presence and third-party corroboration, and any partner promising them on a quarterly timeline is describing something outside their control. Lifewood's AEO services are scoped to the retrieval surface for exactly this reason, with memory-surface movement tracked but never promised.
What does a competent method look like?
A competent ChatGPT visibility method has six components: a fixed prompt set, a pre-work baseline, repeated runs with reported variance, explicit metric definitions, entity work before content work, and an in-house execution capability. Ask the partner to describe their method unprompted and count how many appear.
A fixed prompt set. Typically 20–40 questions in each in-scope language: category questions ("who provides X for enterprises"), comparison questions, and brand questions. Fixed across periods, or nothing is comparable.
A pre-work baseline. Run before any execution starts. Without it, later improvement cannot be attributed to the work — and the partner has removed the only clean evidence of their own value, which is a strange choice to make voluntarily.
Repeated runs and reported variance. Generative answers vary between runs. One run per prompt is an anecdote. Ask for runs-per-prompt and whether variance is reported alongside the mean. The guide to measuring AI visibility without fooling yourself sets out how many runs a prompt set needs before period-to-period movement means anything.
Explicit metric definitions. Share of answer is the proportion of answers to a fixed prompt set in which the brand is mentioned at all. Cited share is the proportion of brand-mentioning answers in which the brand is linked or attributed rather than merely named. Mean rank is the average position the brand takes where the answer returns a ranked list. At minimum:
Share of answer = Answers mentioning the brand ÷ Total answers for the prompt set
Cited share = Answers linking or attributing the brand ÷ Answers mentioning the brand
Mean rank = Average position where the answer returns a ranked list
Three different questions. A partner using them interchangeably is not measuring carefully.
Entity work before content work. Entity work is the job of making a brand resolve as a single, corroborated entity across the web, so that a model can connect every mention of it to the same record. If a model cannot resolve your brand this way, content volume will not fix it. Consistent naming across the estate, declared alternate names and transliterations, resolvable third-party references, and non-contradictory structured data come first; the explainer on entity SEO for AI search covers what that audit looks like in practice.
An execution capability, not a recommendations deck. Ask who writes and publishes. If the answer is "your team, from our brief", price the work you are about to absorb and score them accordingly.
What evidence should you demand before signing?
Demand six artefacts, each of which is cheap for a competent partner to produce and hard to fake convincingly: a raw run file, a before/after pair, a draft prompt set, an entity audit, two published assets, and a written statement of limits.
- A raw run file from a live client period — prompts, timestamps, model and mode, full answer text. Redacted for client identity is fine.
- A before/after pair for one prompt where they moved the outcome, with the dates of the intervening work.
- The prompt set they would use for your category, drafted before the contract. This shows whether they understand how your buyers ask.
- An entity audit of your current estate — the naming inconsistencies, missing corroboration and schema contradictions they can see from outside. A partner who cannot produce a page of specifics here has not looked.
- Two published assets they wrote, with the passage that got cited identified.
- A written statement of what they cannot control. The best partners produce this without being asked.
These six requests sit alongside the broader questions to ask before hiring AEO and GEO help, which cover pricing, team structure and reporting cadence.
How should you score partners?
Score partners on six weighted criteria, with measurement method carrying the most weight at 30% and execution capacity next at 25%, and run a disqualification pass before adding anything up.
| Criterion | Weight | Evidence |
|---|---|---|
| Measurement method | 30% | Baseline procedure, fixed prompt set, runs per prompt, memory/retrieval split, raw file |
| Execution capacity | 25% | In-house writers and reviewers; named counts per language |
| Entity and technical foundation | 15% | Entity audit of your estate; crawl and rendering check |
| Answer-ready content quality | 15% | Published examples; evidence density; liftable passages |
| Reporting and data ownership | 10% | Raw data export; content ownership at contract end |
| Commercial terms | 5% | Notice, ramp, peak capacity |
Disqualifying, at any total score: a guarantee of ChatGPT mentions or rankings; a single blended visibility score with no surface split; no pre-work baseline; refusal to show a raw run file; a prompt set that turns out to be machine-translated for non-English markets.
Weighted score = Σ (criterion score ÷ 5 × weight)
Run the disqualification pass first. A high total score with a disqualifying condition is a well-presented version of the wrong purchase. If you are scoring several candidates at once, the comparison of the best AEO and GEO agencies applies the same criteria to fifteen named providers.
What belongs in the contract?
Five clauses prevent the common disputes: the baseline as a named deliverable, a reporting specification, content ownership, model and method change notification, and a no-guarantee clause stated positively.
- Baseline as a deliverable. Named, dated, delivered before execution begins, with the raw file included.
- Reporting specification. Which metrics, split by surface and by market, at what cadence, with raw data export in a non-proprietary format.
- Content ownership. You own everything produced, including prompt sets and measurement data, and receive it in a usable form at termination.
- Model and method change notification. If the partner changes the model or mode used for measurement, they must say so — otherwise your trend line breaks silently and looks like a performance change.
- No guarantee clause, stated positively. A short paragraph recording that placement in generative answers is not controllable and that the engagement is measured on defined leading indicators. This protects both sides and quietly filters out the partners who will not sign it.
What are the red flags?
The clearest red flags are guaranteed placement, a proprietary score with no formula, volume-first proposals, language coverage counted in tool support, no mention of the entity layer, attribution charts with no confounder discussion, and reluctance to name what will not work.
Guaranteed placement. Nobody controls the output of a model they do not operate.
A proprietary score with no formula. If you cannot reproduce the number from the raw data, it is a marketing device.
Volume-first proposals. "We'll publish sixty articles a quarter" addresses surface area, not citability. The published research points the other way: in the GEO benchmark presented at ACM KDD 2024, tested across 10,000 queries, Quotation Addition raised citation visibility by up to 40% and Statistics Addition by roughly 30%, while keyword stuffing offered little to no gain (about 10% worse than baseline on one Perplexity.ai metric in the same paper).
Language coverage counted in tool support. Ask for in-market writer and reviewer headcount per language instead.
No mention of the entity layer. A partner who goes straight to content has skipped the cheapest available win.
Attribution charts with no confounder discussion. Engines change underneath the measurement. A partner who never mentions this is either not measuring long enough to have noticed, or is choosing not to say.
Reluctance to name what will not work. Every real programme has parts that do not move. A partner who cannot name any is selling certainty rather than method.
How does Lifewood approach ChatGPT visibility?
Lifewood runs ChatGPT visibility inside a single AEO and GEO programme, with the measurement instrument built and operated in-house rather than resold, so memory and retrieval are reported separately by default. A programme that cannot distinguish the two cannot tell a slow win from a failure.
Execution is in-house rather than briefed back to the client. 100+ languages, 40+ delivery centres across 30+ countries and 56,000+ registered contributors mean prompt sets and published content authored by in-market native speakers, including in low-resource languages where most providers fall back to machine translation. Lifewood was founded in 2004 and has over two decades of AI-data delivery behind the programme; engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement.
The retrieval-surface work is delivered as answer-engine optimisation and the longer-horizon narrative work as GEO services, both measured on the same fixed prompt set from the same pre-work baseline. Lifewood's glossary defines share of answer, entity canonicalization and the related terms used in its reporting, so the numbers in a report can be reproduced from the raw run file.