Skip to main content
AEO/GEO

How People Actually Prompt AI Assistants

August 2026 · 8 min read · Updated September 2026

Short answer. Prompt phrasing changes measured brand visibility more than most content changes do. Across 37,804 AI responses from 1,754 prompts on five engines, keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing, ranking-style prompts about 20% higher, and prompt length had effectively zero impact. The question set a vendor chooses to track can move a reported visibility score by more than sixteen points without the brand changing anything at all.

Every AI visibility programme quietly depends on an assumption nobody states: that the prompts being tracked represent how buyers actually ask. Two large 2026 studies tested that assumption from different directions, and both found the choice of question matters more than almost anything the brand does.

This piece sets out what they measured, what follows for building a question set, and why two honest vendors can report very different numbers for the same brand.

Key takeaways

  • Keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing, and ranking-style prompts about 20% higher, across 37,804 AI responses from 1,754 prompts on five engines.
  • Prompt length had effectively zero measured effect on which brands get named, so padding a prompt with extra context is wasted effort.
  • Around 88–92% of human-written prompt pairs cluster above 0.50 cosine similarity, but drifting into the lowest similarity band cut observed visibility by roughly 50%.
  • Mention rate for the same brand ranged from 41.2% on recommendation prompts to 24.5% on research prompts in a separate B2B study, a 16.7-percentage-point spread driven entirely by prompt choice.
  • AI Mode queries average 7.22 words against 4.0 for traditional Google search, so headings phrased as full questions with a direct answer beneath them match the way people actually ask.

What did the prompt-variance study find?

Prompt wording changed measured brand visibility more than any other variable the study tested, while prompt length changed almost nothing.

Ehrlinspiel, Landwehr and Rudzki at Peec AI ran two parallel studies, published 10 June 2026: 288 human-written prompts generating over 17,000 chats, and 54 base prompts expanded into more than 1,000 semantic variations generating over 20,000 chats — across ChatGPT, Gemini, Perplexity, Google AI Mode and Google AI Overviews, spanning 18 sub-verticals in five sectors.

Prompt property Measured effect on brand visibility
Keyword-style phrasing vs conversational Up to +25%
Ranking-style phrasing About +20%
Prompt length Effectively zero
Constraints in the prompt Model-dependent — reduced mentions on ChatGPT and Perplexity, increased them on Gemini and Google AI Overviews

A keyword-style prompt is a short, commercially phrased query built around a category or intent term ("best multilingual data annotation vendor") rather than a full sentence. The interfaces are conversational, so the assumption has been that conversational prompts are the ones that matter. On brand visibility specifically, the opposite holds: a concise commercial phrasing keeps a sharp retrieval anchor, while a conversational or persona-laden prompt broadens the query into educational territory where fewer brands are named.

How far can a prompt drift before the series breaks?

A rewritten prompt below a moderate similarity threshold to the original stops measuring the same thing, even when it asks about the same topic.

Cosine similarity is a 0-to-1 score for how close two pieces of text are in meaning, with 1 meaning identical. Around 88% to 92% of human-written prompt pairs sat above a cosine similarity of 0.50, and about 95% above 0.40. Against a baseline brand mention rate of 4.9%, prompts drifting into the lowest similarity band (0.35–0.39) lost 2.40 percentage points of visibility — roughly a 50% decrease. Stability held above a 0.50 to 0.60 similarity threshold.

Two conclusions follow, and they pull in opposite directions in a useful way.

  • Exhaustive prompt enumeration is unnecessary. Real human phrasings cluster tightly, as the guidance in how to measure AI visibility without fooling yourself also notes — you do not need every possible wording of a question, only one wording inside the cluster.
  • A rewritten prompt is a different measurement. Drifting below roughly 0.50 similarity halves the observed visibility. Question text has to be frozen once a series starts; editing the wording silently changes the baseline, and the trend line becomes unfalsifiable.

Does intent archetype matter more than the engine?

Which type of question a prompt panel favours changes the reported mention rate more than which AI engine it runs on.

A second independent study, from Analyze, measured the same effect from the other direction on B2B prompts specifically: 22,295 AI answers and 115,843 citation events across 460 distinct B2B prompts and 37 organisations on ChatGPT, Perplexity and Google AI Mode.

A prompt archetype is a category of buyer question grouped by intent — recommendation, comparison, or research — rather than by wording. Mention rate varied by archetype from 41.2% for recommendation prompts on Perplexity, to 34.7% for comparison prompts on Google AI Mode, to 24.5% for research prompts on ChatGPT — a spread of 16.7 percentage points. Within engines the spread was still 12.2 points on Perplexity, 9.3 on Google AI Mode and 8.5 on ChatGPT.

A panel dominated by recommendation prompts, measured mainly on Perplexity, lands near the top of the range. The same organisation measured on research prompts on ChatGPT lands near the bottom — not through dishonesty, simply by choosing which kinds of questions to track, an effect the share of answer metric is built to expose rather than hide.

How do you build a question set that is not self-flattering?

A defensible prompt set fixes its mix of question types before measurement starts and reports each type separately instead of blending them into one number.

Mention rate is the share of tracked AI answers in which a given brand is named at all, regardless of ranking or sentiment.

Archetype Example Typical mention rate Use it to measure
Recommendation / shortlist "What are the best X providers for enterprise?" Highest Whether you are in the consideration set
Comparison / alternatives "X vs Y for a multilingual programme" Middle How you are positioned against named rivals
Research / how-to "How does X actually work?" Lowest Whether your explanatory content is retrievable
Accuracy "What does the company do?" Not a visibility metric Whether what is said about you is true
  1. Include all archetypes deliberately, and report them separately. A blended number is dominated by whichever archetype you happened to include most.
  2. Fix the proportions before you start. If the mix changes between reporting periods, the trend line is an artefact of the mix rather than a measurement of anything.
  3. Prefer concise, commercially phrased questions for visibility measurement, since that produces the sharp retrieval anchor — but do not then claim the result represents all user behaviour.
  4. Freeze the wording. Below roughly 0.50 similarity you are measuring a different question.
  5. Do not pad prompts. Length has no measured effect, so extra context only makes the question harder to reproduce.
  6. Ask a vendor for the archetype mix before comparing their number to anyone else's, using the framework in an AI visibility audit. Two honest measurements of the same brand can differ by sixteen points on mix alone.

This is the most under-examined lever in the category. Everyone argues about which tool to buy; almost nobody asks what proportion of the tracked prompts are recommendation prompts. The second question determines the number far more than the first.

What do people actually type?

Queries reaching AI answer surfaces are longer and more fully formed than search-engine queries, even though length itself does not change brand visibility.

Semrush data compiled by AEO Vision puts average Google AI Mode query length at 7.22 words against 4.0 for traditional Google search. That is a writing instruction rather than a measurement instruction: a 7.22-word query is a question, not a keyword fragment. Pages whose headings are phrased as full questions with a direct answer beneath them match those queries directly; pages organised around keyword targets do not.

What are the limits of this research?

Both studies are recent, well-designed, and narrower than the headline numbers might suggest, so they should inform prompt design rather than settle it outright.

  • The Peec AI study is recent and single-team. It is unusually large and well-designed for this field, and it has not yet been independently replicated.
  • The archetype study is B2B-specific. 460 prompts and 37 organisations is a solid sample but a narrow domain; consumer categories may behave differently.
  • Mention rate is not accuracy. Every figure here counts whether a brand was named, not whether the sentence about it was true.
  • These are relative effects, not levers a brand controls directly. Choosing better prompts changes what you measure, not what buyers ask.

How does Lifewood apply this?

Lifewood treats the prompt registry itself as the most consequential artefact in a visibility programme, ahead of the dashboard that reads it.

Archetype proportions are fixed before the first run and reported separately rather than blended, so a movement can be traced to a category of question rather than absorbed into one number, in line with the approach described for measuring LLM visibility across engines. Wording is frozen at the start of a series and versioned when it changes, with the change logged, because the evidence says a rewritten prompt is a different measurement rather than a refined one. Prompts are kept concise and commercially phrased for visibility measurement, and that limitation is stated in the report rather than left implied.

For non-English markets the registry is authored natively rather than translated, since a translated list measures the translation rather than the market. Lifewood's 100+ languages and 40+ delivery centres across 30+ countries are what make native authorship a staffing decision rather than a compromise, and the same discipline applies to GEO service work more broadly.

Frequently asked questions

Substantially. Across 37,804 responses, keyword-style prompts produced up to 25% higher average brand visibility than conversational phrasing and ranking-style prompts about 20% higher, while prompt length had effectively no impact. Phrasing is one of the largest single variables in any AI visibility measurement.

Fewer than most tools imply. Around 88–92% of human-written prompt pairs cluster above 0.50 cosine similarity, so one phrasing inside that cluster represents the group. What matters more is freezing the wording, because drifting below roughly 0.50 similarity cut observed visibility by about half.

Frequently because of prompt mix rather than measurement error. Mention rates ranged from 41.2% for recommendation prompts on Perplexity to 24.5% for research prompts on ChatGPT — a 16.7-point spread. A panel weighted toward recommendation prompts reports a much higher number for the same brand.

Recommendation and shortlist prompts, followed by comparison and alternatives, with research and how-to prompts lowest. The spread within a single engine was 8.5 to 12.2 percentage points, so archetype mix matters even when the engine is held constant.

Match the question shape, not the length. AI Mode queries average 7.22 words against 4.0 for traditional search, so headings phrased as full questions with a direct answer beneath them match well. Prompt length itself had no measurable effect on which brands get named, so padding is wasted effort.

Ask for the prompt list, the archetype proportions, and confirmation that the wording is frozen. Then require results reported per archetype and per engine rather than blended. Sixteen points of movement are available through mix selection alone.

Sources and further reading

  1. Prompt Tracking: Does prompt variance % impact brand mentions? (Search Engine Journal, on the Peec AI study)
  2. Prompt Type Changes AI Mention Rate 17 Points (Analyze, State of AI search: prompt archetypes)
  3. Google AI Mode GEO Statistics 2026 (AEO Vision, citing Semrush query-length data)

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team