Skip to main content
AEO/GEO

How Do Voice Assistants Pick Their One Answer?

August 2026 · 7 min read · Updated September 2026

Short answer. By extraction, not ranking. The assistant treats your spoken question as a search, then looks for one passage short enough to read aloud in a few seconds — usually the featured snippet at position zero, sometimes a source cited inside an AI Overview. One widely cited analysis of 10,000 voice searches found roughly 40% of answers came from featured snippets, and a snippet-holding page is far more likely to be read aloud than one ranking lower without it. For "near me" and "open now" questions, the answer comes from the Google Business Profile instead.

Voice search does not return a list. It returns a sentence, and then it stops. Everything that makes typed search forgiving — ten results, a page two, a user willing to scroll — is absent, which is why the selection mechanics matter more here than anywhere else in search. This piece covers why the channel is winner-take-all, which sources the spoken answer is drawn from, what makes one passage speakable and another unusable, and the order in which to fix things.

Key takeaways

  • Voice search is winner-take-all: the assistant reads one answer and stops, so ranking 2–10 without the featured snippet returns nothing at all.
  • Roughly 40% of voice answers are read from the featured snippet at position zero, and a snippet-holding page is far more likely to be voiced than one ranking lower.
  • For "near me" and "open now" questions, the spoken answer is pulled almost entirely from the Google Business Profile rather than from a website.
  • A voice-ready answer is roughly 23–41 words long, sits within the first 30–60 words of its section, and is phrased the way the question would actually be spoken.
  • Voice assistants perform worse in lower-resource languages, so an answer that exists only in English is effectively invisible to spoken queries in every other market.

Why is voice winner-take-all?

Because there is no page two. The assistant reads one answer and stops, so second place returns nothing at all.

Type a query and you get ten links to choose from. Ask it aloud and you get one response, spoken. A featured snippet — the boxed answer Google pulls above the normal results for a query — is what most voice answers are read from, so winning it is worth more than ranking 2–10 combined for the same question, because those lower positions are never voiced. Voice now accounts for roughly 31% of searches and voice commerce is estimated at around $86 billion in 2025, heading toward $164 billion by 2028.

The queries themselves are different too. Spoken questions run noticeably longer than typed searches and arrive as complete sentences: "where's the best pizza place near me that's open right now?" rather than "pizza NYC". They also skew heavily toward immediate and local intent, which is why the stakes are so concrete — a local voice search often converts to an in-store visit within a day. Optimising only for short typed keywords makes a business effectively invisible to this entire channel.

Source Share of voice answers Query type it serves Verdict
Featured snippet ~40.7% of answers in Backlinko's 10,000-search analysis; snippet pages are far more likely to be read aloud Informational: what is, how much, how do I The primary source
Top organic results (no snippet) A meaningful minority of remaining answers Questions with no clean snippet available The fallback
Google Business Profile Near-exclusive for local intent "Near me", "open now", "closest" Owns local answers
AI Overview citations Growing through 2026 Complex or multi-part questions The rising channel

Estimates vary by study — some analyses put featured-snippet sourcing even higher. All agree position zero dominates.

Where does the spoken answer actually come from?

From one of four places, depending on what was asked — and each rewards a different asset.

For informational questions, the assistant reads the featured snippet, which is why snippet ownership, not average ranking position, is the metric that matters: pages already ranking near the top for a question query have the best chance of earning it. For local questions with near-me or open-now intent, Google bypasses websites almost entirely and reads from the Google Business Profile — a business's free, editable Google listing showing name, hours, services and ratings. For more complex questions, assistants increasingly read from sources cited inside AI Overviews — the AI-generated summary Google places above standard search results. And when no clean snippet exists, they fall back to the top few organic results.

Platform quality differences are real but narrowing. Google Assistant leads recent benchmarks with roughly 93.7% query comprehension and an 87.4% correct-answer rate, ahead of Siri and Alexa. The practical implication is that the pipeline is shared: voice answers are pulled from the same organic results, snippets and profiles that power ordinary search, so improving snippet-readiness lifts both surfaces at once.

What makes one passage readable and another unusable?

Length, position and phrasing. The assistant needs a self-contained answer it can speak in a breath.

What gets read aloud What cannot be spoken
A direct answer in the first 30–60 words Answers buried below preamble
Answers of roughly 23–41 words — one spoken breath Facts that only exist in a table image or PDF
Headings that mirror the spoken question Keyword-fragment headings nobody says aloud
Natural, active phrasing: "you can do this by…" Long paragraphs with no self-contained sentence
Pages that load quickly Content in one language when the question is asked in another

The page is written for the reader; the extractable passage is what the assistant speaks. If it cannot be lifted cleanly, it will not be voiced.

The structure that consistently works is simple: question heading, direct answer paragraph, then expanded explanation and supporting detail beneath. It serves the assistant, which extracts the top passage, and the human, who wants the depth underneath. Speed matters as a gate rather than a tiebreaker — voice results tend to load noticeably faster than non-voice results, and slow pages are filtered out before their content is ever considered. This is the same passage-level discipline covered in how people actually prompt AI assistants, and it applies just as much to typed prompts as to spoken ones.

The multilingual dimension is where most brands quietly lose. A spoken question in Bahasa Indonesia, Arabic or Hindi is matched against content in that language, and assistants perform noticeably worse in lower-resource languages and regional accents — a pattern also visible in the English-language bias running through AI search more broadly. If your answer blocks, profile content and FAQs exist only in English, you are absent from every spoken answer in every other market, and nothing in your analytics reports it. Producing natural, native-speaker answer content across languages is exactly the kind of delivery work an AEO and GEO practice takes on, and here it is a visibility investment rather than a translation task.

What should you fix first?

In this order, because the effort-to-effect ratio differs sharply.

  1. Find the questions you already nearly own. Pages ranking near the top for a question query are the realistic snippet targets. Start there rather than with your hardest keywords.
  2. Put the answer in the first sentence. Question in the H2 or H3, direct answer immediately beneath in 23–41 words, detail afterwards. No preamble above the answer. This is the same question-heading, answer-first structure that answer engines reward generally, not just for voice.
  3. Write headings the way people speak. "How much does a dental cleaning cost in Bangkok?" not "Affordable dental pricing". Voice queries run long, so target full spoken sentences.
  4. Complete your Google Business Profile if any of your queries are local. Hours, services, ratings and address are read directly from it, and no amount of website work substitutes — the same profile-first logic covered in how local businesses show up in AI search.
  5. Get the page loading quickly. Voice results load faster than average; speed is a prerequisite, not an optimisation.
  6. Move facts out of images and PDFs into HTML. A price or statistic that exists only inside a graphic cannot be spoken, and it is invisible to how Google AI Overviews chooses what to cite for the same reason.
  7. Do all of it in every language you sell in. Answer blocks and profile content are language-specific assets; coverage in one buys you nothing in another. Multilingual GEO services exist specifically to close this gap at scale.

The first three items are the same structural work a managed AEO and GEO practice applies across a site: question-shaped headings, a self-contained answer beneath each one, evidence below that.

Frequently asked questions

No. Voice answers are pulled from the same organic results, featured snippets and business profiles that power typed search — there is no separate "voice index." That is why improving a page's snippet-readiness lifts its visibility in both typed and spoken search at once, rather than requiring separate work for each channel.

Roughly 23–41 words — short enough to be spoken in one breath, self-contained enough to make sense without the surrounding page. It should also sit within the first 30–60 words of its section, directly beneath a question-shaped heading, with supporting detail placed afterward rather than before it.

Structured data still helps engines map what your content answers, but FAQ rich results stopped appearing in standard Google Search in May 2026. Well-organised question-and-answer content remains valuable regardless of the markup, both for voice extraction and for citation by AI answer engines.

Prioritise the Google Business Profile over everything else. For near-me and open-now questions, that profile is effectively the only source the assistant reads from — name, hours, services and ratings are read straight from it, and website content plays little to no role in that decision.

Because only one answer is spoken. A page holding the featured snippet is far more likely to be read aloud than a page ranking lower without one, so the gap between position zero and position two is not incremental — for voice, it is total.

Sources and further reading

  1. Digital Applied, "Voice Search Statistics 2026: 100+ Data Points and Trends" — Backlinko's 10,000-search analysis and the 40.7% featured-snippet share.
  2. Digital Applied, "Voice Search SEO 2026: Optimize for 8.4 Billion Devices" — winner-take-all dynamics, query length and page-speed considerations for voice.
  3. SEO Scale Up, "Voice Search Statistics 2026" — Google Assistant's comprehension and correct-answer benchmarks, and voice commerce sizing.
  4. Single Grain, "Best Voice Search Optimization Tools in 2026" — assistants selecting a single spoken answer from snippets or AI Overview citations.
  5. BigFin SEO, "What Is Voice Search Optimization? A 2026 Guide" — the 23–41 word answer-length guidance and voice-versus-text query structure.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team