LIFEWOOD
Ready100
AEO/GEO

How People Actually Prompt AI Assistants

Short answer. Prompt phrasing changes measured brand visibility more than most content changes do. Across 37,804 AI responses from 1,754 prompts on five engines, keyword-style prompts…

Lifewood Data Technology · August 2026 · 6 min read

Download PDF

Short answer. Prompt phrasing changes measured brand visibility more than most content changes do. Across 37,804 AI responses from 1,754 prompts on five engines, keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing, ranking-style prompts about 20% higher, and prompt length had effectively zero impact. Which means the question set you choose largely determines the number you report — a vendor can move your AI visibility by more than sixteen points without touching your website.

Every AI visibility programme quietly depends on an assumption nobody states: that the prompts being tracked represent how buyers actually ask. Two large 2026 studies tested that assumption from different directions, and both found the choice of question matters more than almost anything the brand does.

This piece sets out what they measured, what follows for building a question set, and why two honest vendors can report very different numbers for the same brand.


What did the prompt-variance study find?

Ehrlinspiel, Landwehr and Rudzki at Peec AI ran two parallel studies, published 10 June 2026: 288 human-written prompts generating over 17,000 chats, and 54 base prompts expanded into more than 1,000 semantic variations generating over 20,000 chats — across ChatGPT, Gemini, Perplexity, Google AI Mode and Google AI Overviews, spanning 18 sub-verticals in five sectors.

Prompt property Measured effect on brand visibility
Keyword-style phrasing vs conversational Up to +25%
Ranking-style phrasing About +20%
Prompt length Effectively zero
Constraints in the prompt Model-dependent — reduced mentions on ChatGPT and Perplexity, increased them on Gemini and Google AI Overviews

The counter-intuitive headline is the first row. The interfaces are conversational, so the assumption has been that conversational prompts are the ones that matter. On brand visibility specifically, the opposite holds: a concise commercial phrasing keeps a sharp retrieval anchor, while a conversational or persona-laden prompt broadens the query into educational territory where fewer brands are named.

Length does nothing. Phrasing does a great deal. Those two findings together are the whole practical lesson.


How far can a prompt drift before the series breaks?

The same study measured stability, and this is the more useful half of it.

Around 88% to 92% of human-written prompt pairs sat above a cosine similarity of 0.50, and about 95% above 0.40. Against a baseline brand mention rate of 4.9%, prompts drifting into the lowest similarity band (0.35–0.39) lost 2.40 percentage points of visibility — roughly a 50% decrease. Stability held above a 0.50 to 0.60 similarity threshold.

Two conclusions follow, and they pull in opposite directions in a useful way.

  • Exhaustive prompt enumeration is unnecessary. Real human phrasings cluster tightly. You do not need every possible wording of a question; you need one wording inside the cluster.
  • But a rewritten prompt is a different measurement. Drifting below roughly 0.50 similarity halves the observed visibility. That is why question text has to be frozen once a series starts. Editing the wording silently changes the baseline, and the trend line becomes unfalsifiable.

Does intent archetype matter more than the engine?

A second independent study measured the same effect from the other direction, on B2B prompts specifically. Analyze tracked 22,295 AI answers and 115,843 citation events across 460 distinct B2B prompts and 37 organisations on ChatGPT, Perplexity and Google AI Mode.

Mention rate varied by archetype from 41.2% for recommendation prompts on Perplexity, to 34.7% for comparison prompts on Google AI Mode, to 24.5% for research prompts on ChatGPT — a spread of 16.7 percentage points. Within engines the spread was still 12.2 points on Perplexity, 9.3 on Google AI Mode and 8.5 on ChatGPT.

The implication is worth stating plainly. A panel dominated by recommendation prompts, measured mainly on Perplexity, lands near the top of the range. The same organisation measured on research prompts on ChatGPT lands near the bottom. A vendor can move your reported AI visibility by more than sixteen points without touching your website — not through dishonesty, simply by choosing which kinds of questions to track.


How do you build a question set that is not self-flattering?

Archetype Example Typical mention rate Use it to measure
Recommendation / shortlist "What are the best X providers for enterprise?" Highest Whether you are in the consideration set
Comparison / alternatives "X vs Y for a multilingual programme" Middle How you are positioned against named rivals
Research / how-to "How does X actually work?" Lowest Whether your explanatory content is retrievable
Accuracy "What does [company] do?" Not a visibility metric Whether what is said about you is true
  1. Include all archetypes deliberately, and report them separately. A blended number is dominated by whichever archetype you happened to include most.
  2. Fix the proportions before you start. If the mix changes between reporting periods, the trend line is an artefact of the mix rather than a measurement of anything.
  3. Prefer concise, commercially-phrased questions for visibility measurement, since that produces the sharp retrieval anchor — but do not then claim the result represents all user behaviour.
  4. Freeze the wording. Below roughly 0.50 similarity you are measuring a different question.
  5. Do not pad prompts. Length has no measured effect, so extra context only makes the question harder to reproduce.
  6. Ask a vendor for the archetype mix before comparing their number to anyone else's. Two honest measurements of the same brand can differ by sixteen points on mix alone.

This is the most under-examined lever in the category. Everyone argues about which tool to buy; almost nobody asks what proportion of the tracked prompts are recommendation prompts. The second question determines the number far more than the first.


What do people actually type?

The queries reaching these surfaces are longer and more fully formed than search queries, even though length does not change brand visibility. Semrush data compiled by AEO Vision puts average Google AI Mode query length at 7.22 words against 4.0 for traditional Google search.

That is a writing instruction rather than a measurement instruction. A 7.22-word query is a question, not a keyword fragment. Pages whose headings are phrased as those questions, and answered directly in the first sentence beneath, match them directly. Pages organised around keyword targets do not.


Limits worth stating

  • The Peec AI study is recent and single-team. It is unusually large and well-designed for this field, and it has not yet been independently replicated.
  • The archetype study is B2B-specific. 460 prompts and 37 organisations is a solid sample but a narrow domain; consumer categories may behave differently.
  • Mention rate is not accuracy. Every figure here counts whether a brand was named, not whether the sentence about it was true.
  • These are relative effects, not levers you control. Choosing better prompts changes what you measure, not what buyers ask.

How Lifewood approaches this

Lifewood treats the prompt registry as the most consequential artefact in a programme, not the dashboard that reads it. Archetype proportions are fixed before the first run and reported separately rather than blended, so a movement can be traced to a category of question rather than absorbed into one number.

Wording is frozen at the start of a series and versioned when it changes, with the change logged, because the evidence says a rewritten prompt is a different measurement rather than a refined one. Prompts are kept concise and commercially phrased for visibility measurement, and that limitation is stated in the report rather than left implied.

For non-English markets the registry is authored natively rather than translated, since a translated list measures the translation. 50+ languages and 40+ delivery centres across 30+ countries are what make that a staffing decision rather than a compromise. See how to measure AI visibility without fooling yourself and what AI visibility tools can and cannot measure.


Sources and further reading

  • Ehrlinspiel, Landwehr & Rudzki (Peec AI), prompt variance study, SSRN, published 10 June 2026: 37,804 AI responses from 1,754 prompts across five engines. Reported by Search Engine Journal.
  • Analyze, State of AI search: prompt archetypes: 22,295 answers, 115,843 citation events, 460 B2B prompts, 37 organisations.
  • Semrush AI Mode query-length data, compiled by AEO Vision.

Frequently asked questions

Substantially. Across 37,804 responses, keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing and ranking-style prompts about 20% higher, while prompt length had effectively no impact. Phrasing is one of the largest single variables in any AI visibility measurement.

Fewer than most tools imply. Around 88–92% of human-written prompt pairs cluster above 0.50 cosine similarity, so one phrasing inside that cluster represents the group. What matters more is freezing the wording, because drifting below roughly 0.50 similarity cut observed visibility by about half.

Frequently because of prompt mix rather than measurement error. Mention rates ranged from 41.2% for recommendation prompts on Perplexity to 24.5% for research prompts on ChatGPT — a 16.7-point spread. A panel weighted toward recommendation prompts reports a much higher number for the same brand.

Recommendation and shortlist prompts, followed by comparison and alternatives, with research and how-to prompts lowest. The spread within a single engine was 8.5 to 12.2 percentage points, so archetype mix matters even when the engine is held constant.

Match the question shape, not the length. AI Mode queries average 7.22 words against 4.0 for traditional search, so headings phrased as full questions with a direct answer beneath them match well. But prompt length itself had no measurable effect on which brands get named, so padding is wasted effort.

Ask for the prompt list, the archetype proportions, and confirmation that the wording is frozen. Then require results reported per archetype and per engine rather than blended. Sixteen points of movement are available through mix selection alone.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team