Short answer. By extraction, not ranking. The assistant treats your spoken question as a search, then looks for one passage short enough to read aloud in a few seconds — usually the featured snippet at position zero, sometimes a source cited inside an AI Overview. One widely cited analysis of 10,000 voice searches found roughly 40% of answers came from featured snippets, and a snippet-holding page is far more likely to be read aloud than one ranking lower without it. For "near me" and "open now" questions, the answer comes from the Google Business Profile instead.
Voice search does not return a list. It returns a sentence, and then it stops. Everything that makes typed search forgiving — ten results, a page two, a user willing to scroll — is absent, which is why the selection mechanics matter more here than anywhere else in search. This piece covers why the channel is winner-take-all, which sources the spoken answer is drawn from, what makes one passage speakable and another unusable, and the order in which to fix things.
Key takeaways
- Voice search is winner-take-all: the assistant reads one answer and stops, so ranking 2–10 without the featured snippet returns nothing at all.
- Roughly 40% of voice answers are read from the featured snippet at position zero, and a snippet-holding page is far more likely to be voiced than one ranking lower.
- For "near me" and "open now" questions, the spoken answer is pulled almost entirely from the Google Business Profile rather than from a website.
- A voice-ready answer is roughly 23–41 words long, sits within the first 30–60 words of its section, and is phrased the way the question would actually be spoken.
- Voice assistants perform worse in lower-resource languages, so an answer that exists only in English is effectively invisible to spoken queries in every other market.
Why is voice winner-take-all?
Because there is no page two. The assistant reads one answer and stops, so second place returns nothing at all.
Type a query and you get ten links to choose from. Ask it aloud and you get one response, spoken. A featured snippet — the boxed answer Google pulls above the normal results for a query — is what most voice answers are read from, so winning it is worth more than ranking 2–10 combined for the same question, because those lower positions are never voiced. Voice now accounts for roughly 31% of searches and voice commerce is estimated at around $86 billion in 2025, heading toward $164 billion by 2028.
The queries themselves are different too. Spoken questions run noticeably longer than typed searches and arrive as complete sentences: "where's the best pizza place near me that's open right now?" rather than "pizza NYC". They also skew heavily toward immediate and local intent, which is why the stakes are so concrete — a local voice search often converts to an in-store visit within a day. Optimising only for short typed keywords makes a business effectively invisible to this entire channel.
| Source | Share of voice answers | Query type it serves | Verdict |
|---|---|---|---|
| Featured snippet | ~40.7% of answers in Backlinko's 10,000-search analysis; snippet pages are far more likely to be read aloud | Informational: what is, how much, how do I | The primary source |
| Top organic results (no snippet) | A meaningful minority of remaining answers | Questions with no clean snippet available | The fallback |
| Google Business Profile | Near-exclusive for local intent | "Near me", "open now", "closest" | Owns local answers |
| AI Overview citations | Growing through 2026 | Complex or multi-part questions | The rising channel |
Estimates vary by study — some analyses put featured-snippet sourcing even higher. All agree position zero dominates.
Where does the spoken answer actually come from?
From one of four places, depending on what was asked — and each rewards a different asset.
For informational questions, the assistant reads the featured snippet, which is why snippet ownership, not average ranking position, is the metric that matters: pages already ranking near the top for a question query have the best chance of earning it. For local questions with near-me or open-now intent, Google bypasses websites almost entirely and reads from the Google Business Profile — a business's free, editable Google listing showing name, hours, services and ratings. For more complex questions, assistants increasingly read from sources cited inside AI Overviews — the AI-generated summary Google places above standard search results. And when no clean snippet exists, they fall back to the top few organic results.
Platform quality differences are real but narrowing. Google Assistant leads recent benchmarks with roughly 93.7% query comprehension and an 87.4% correct-answer rate, ahead of Siri and Alexa. The practical implication is that the pipeline is shared: voice answers are pulled from the same organic results, snippets and profiles that power ordinary search, so improving snippet-readiness lifts both surfaces at once.
What makes one passage readable and another unusable?
Length, position and phrasing. The assistant needs a self-contained answer it can speak in a breath.
| What gets read aloud | What cannot be spoken |
|---|---|
| A direct answer in the first 30–60 words | Answers buried below preamble |
| Answers of roughly 23–41 words — one spoken breath | Facts that only exist in a table image or PDF |
| Headings that mirror the spoken question | Keyword-fragment headings nobody says aloud |
| Natural, active phrasing: "you can do this by…" | Long paragraphs with no self-contained sentence |
| Pages that load quickly | Content in one language when the question is asked in another |
The page is written for the reader; the extractable passage is what the assistant speaks. If it cannot be lifted cleanly, it will not be voiced.
The structure that consistently works is simple: question heading, direct answer paragraph, then expanded explanation and supporting detail beneath. It serves the assistant, which extracts the top passage, and the human, who wants the depth underneath. Speed matters as a gate rather than a tiebreaker — voice results tend to load noticeably faster than non-voice results, and slow pages are filtered out before their content is ever considered. This is the same passage-level discipline covered in how people actually prompt AI assistants, and it applies just as much to typed prompts as to spoken ones.
The multilingual dimension is where most brands quietly lose. A spoken question in Bahasa Indonesia, Arabic or Hindi is matched against content in that language, and assistants perform noticeably worse in lower-resource languages and regional accents — a pattern also visible in the English-language bias running through AI search more broadly. If your answer blocks, profile content and FAQs exist only in English, you are absent from every spoken answer in every other market, and nothing in your analytics reports it. Producing natural, native-speaker answer content across languages is exactly the kind of delivery work an AEO and GEO practice takes on, and here it is a visibility investment rather than a translation task.
What should you fix first?
In this order, because the effort-to-effect ratio differs sharply.
- Find the questions you already nearly own. Pages ranking near the top for a question query are the realistic snippet targets. Start there rather than with your hardest keywords.
- Put the answer in the first sentence. Question in the H2 or H3, direct answer immediately beneath in 23–41 words, detail afterwards. No preamble above the answer. This is the same question-heading, answer-first structure that answer engines reward generally, not just for voice.
- Write headings the way people speak. "How much does a dental cleaning cost in Bangkok?" not "Affordable dental pricing". Voice queries run long, so target full spoken sentences.
- Complete your Google Business Profile if any of your queries are local. Hours, services, ratings and address are read directly from it, and no amount of website work substitutes — the same profile-first logic covered in how local businesses show up in AI search.
- Get the page loading quickly. Voice results load faster than average; speed is a prerequisite, not an optimisation.
- Move facts out of images and PDFs into HTML. A price or statistic that exists only inside a graphic cannot be spoken, and it is invisible to how Google AI Overviews chooses what to cite for the same reason.
- Do all of it in every language you sell in. Answer blocks and profile content are language-specific assets; coverage in one buys you nothing in another. Multilingual GEO services exist specifically to close this gap at scale.
The first three items are the same structural work a managed AEO and GEO practice applies across a site: question-shaped headings, a self-contained answer beneath each one, evidence below that.