Skip to main content
AEO/GEO

Technical SEO for AI Search Visibility

July 2026 · 8 min read · Updated September 2026

Short answer. AI search visibility depends on the same technical foundation as ordinary search: crawlable URLs, main content present in the served HTML, correct status codes, self-consistent canonicals, real internal links, an accurate sitemap, and pages that render without a browser. Generative features run on the same crawling and indexing systems as ranked search, so a page an engine cannot fetch and parse is invisible to both — and this failure is undetectable from inside a browser, where everything looks fine.

Key takeaways

  • Retrieval is a hard filter: a page that is blocked, orphaned, noindexed, or misrendered never enters the candidate pool an AI answer can cite from.
  • Checks have a strict order — robots rules, status codes, canonicals, discoverability — because a failure at an early step makes later content diagnostics meaningless.
  • Heavy client-side rendering is the single most common cause of an otherwise good page being invisible to non-browser fetchers.
  • Bot-management and CDN defaults can silently serve AI agents a shorter page or a block page, a failure only a byte-count comparison across user agents reveals.
  • Page speed and Core Web Vitals are not a citation factor on their own, but the rendering work behind slow pages often causes the same content to be missing from the served HTML.

Why does crawlability come before anything else?

Retrieval acts as a filter with a hard edge, and nothing downstream of it matters until a page passes through. Crawlability is a page's basic reachability to an automated fetcher — whether it is allowed, discoverable, and returns usable content and status codes. A page that is blocked, orphaned, noindexed, returning the wrong status code, or hidden behind a rendering path a fetcher cannot follow never enters the candidate pool. Evidence density, passage structure and schema have no effect on a page that was never a candidate in the first place.

The order of checks matters, because a failure at an early step makes later diagnostics meaningless:

  1. robots.txt — is the path allowed, for every agent that matters, not only for a wildcard?
  2. Robots meta and X-Robots-Tag — is the page noindex by accident? Staging directives shipped to production are a routine cause.
  3. Status codes — 200 for content, single-hop 301 for moves, and no soft 404s returning 200 with an error page.
  4. Canonical — self-referential on the canonical URL, and consistent with what the sitemap and internal links point at.
  5. Discoverability — is the page linked from another crawlable page, or reachable only through a search box or a JavaScript filter?

Only after all five pass is a content diagnosis worth running. Many of the same defects covered in this technical AEO checklist show up first as a crawlability failure rather than a content gap.

What should JavaScript sites verify?

The served HTML — the markup a fetcher receives before any client-side script runs — is what decides whether content exists for a non-browser agent, and checking it in a browser hides the problem because a browser executes the JavaScript that a fetcher may not.

The test is simple: load the page with JavaScript disabled, or fetch the raw HTML directly. What remains is approximately what a fetcher that does not execute JavaScript receives. In that raw HTML, look for the main content in full rather than a loading shell or placeholder; real anchors with href attributes, since navigation built only from click handlers hides the link graph that carries authority to deep pages; title, canonical and meta directives, which can arrive missing or wrong when injected at runtime; numbers that are actually numbers rather than counters that animate up from zero and render as "0" in the served HTML; and content behind tabs and accordions, where collapsed is fine but conditionally rendered is not, because a component that mounts a panel's contents only when opened serves a fetcher a different answer than it serves a human.

Where rendering is required to see the content, server-side rendering or pre-rendering is the fix. It is not a visibility guarantee — relevance, evidence and competition still decide the outcome — but it removes a class of failure no content work can compensate for.

Are the AI agents being served the same page you are?

No page check is complete until it accounts for the possibility that a specific agent receives less than a browser does. A page can pass every crawlability and rendering check for a browser and still arrive truncated at a given AI user agent, because rate limiters, bot-management rules and CDN defaults commonly return a shorter page, a challenge page, or a 403 to agents they do not recognise — and none of that is visible from inside the site.

The verification is a byte-count comparison: bytes served to the agent in question, divided by bytes served to a browser fetch of the same URL. Run it per relevant user agent against a handful of important pages; anything materially below 1.0 is a delivery defect, not a content one. Two points are worth flagging separately. Training crawlers and retrieval crawlers are different bots doing different jobs, and common defaults block them under one toggle — blocking a retrieval fetcher makes citation impossible, a distinction covered in more depth in this guide to AI crawlers and AI search visibility. And a managed bot rule a team did not configure is still that team's rule: platform and CDN defaults change without a deploy on the publisher's side, so this check belongs on a schedule rather than a one-time launch item.

Do page speed and Core Web Vitals matter here?

Page speed matters for users and for ordinary search quality, but it is not a citation switch on its own. There is no evidence that a faster page is more likely to be quoted by an AI answer, so treating performance work as AEO work misallocates effort.

The version of this that does matter is the shared cause behind both problems: heavy client-side rendering slows a page and removes content from the served HTML at the same time, and layout that injects the answer late hurts both outcomes. Optimising images, scripts and fonts is worth doing on its own merits, and stripping useful content to hit a performance score is not. The one performance-adjacent item with a direct effect on citability is crawl efficiency — a site that times out under a fetcher's rate limit gets less of itself fetched, regardless of how fast it loads for a human visitor.

What should a technical AI-search audit include?

A technical AI-search audit is a template-level review that checks whether a site's delivery mechanics let content be fetched, parsed and indexed at all. Run these checks as a set, on templates rather than on individual pages, because one template defect affects every page built from it:

  • Index coverage and the reasons for exclusion
  • Served-HTML completeness on each page template
  • Byte-count parity across relevant user agents
  • Canonical consistency between page, sitemap and internal links
  • Redirect chains reduced to single hops
  • Status codes, including soft 404s
  • Internal-link depth and orphan detection
  • Sitemap accuracy — canonical, indexable URLs only, with real lastmod values
  • Structured data that matches the visible page and does not contradict it
  • Titles and descriptions present in the served HTML
  • Publication and modification dates that reflect actual changes

Fixing template-level defects before rewriting articles is the more efficient order: one broken template is cheaper to repair than forty rewritten pages, and it is more often the actual cause of a citation gap than the writing described in what actually gets you cited by AI answer engines.

How does Lifewood sequence this work?

Lifewood treats delivery as a gate that runs before content work, on client sites and its own, rather than as one item on a longer list. The sequence is fixed: entity resolution, then machine readability, then publishing, and pages are checked as served rather than as rendered — JavaScript disabled, byte counts compared per agent, tab and accordion contents confirmed present in the markup — because this is the failure that silently invalidates every content decision made downstream of it.

The reason the order is non-negotiable is measurement. Publishing into an unresolved delivery defect produces no movement and no way to tell whether the content itself was the problem. Establishing a baseline, fixing delivery, publishing, and re-measuring is what keeps a result attributable to a specific change rather than lost in an unmeasured backlog of defects, an approach detailed further in Lifewood's AEO and GEO service overviews.

Frequently asked questions

Yes, with the caveats about added complexity and delay that Google's own documentation sets out. The larger issue is that not every fetcher that matters renders JavaScript at all. Content that exists only after hydration is available to some systems and absent for others, which is why the served-HTML check is worth running regardless of what any single engine can do.

No. It removes one specific and common failure — content absent from the served HTML — and nothing more. Relevance, evidence, entity clarity and competition still decide whether a passage gets used once it is retrievable. Treat server-side rendering as a prerequisite fix, not a ranking lever.

Fetch the page as the agent in question and compare the byte count and visible text against a browser fetch of the same URL. If the agent receives materially less, it is a delivery problem and no rewrite will fix it. If the agent receives the full page and the passage still goes unused, the problem is the content itself.

Every canonical page intended for indexing, yes, with an honest `lastmod` and no non-canonical, redirected or noindexed URLs included. A sitemap containing URLs that redirect or error reduces the trust placed in the whole file, and a `lastmod` that updates on every deploy tells a crawler nothing useful.

Collapsed is fine; conditionally rendered is not. If the answer panel's text exists in the served HTML and is merely hidden by CSS, it is present and citable. If the component mounts the panel's contents only when a user clicks, a crawler receives the questions and just one answer — a common and quiet defect on otherwise well-built pages.

It is a low-cost addition, not a route into any engine's answers — Google Search does not use it, and no major engine treats it as a requirement. The fuller reasoning is in this piece on [what llms.txt actually is](/blogs/what-is-llms-txt); treat it as optional housekeeping after the fundamentals above, never as a substitute for them.

Sources and further reading

  1. Google Search Central: Optimizing your website for generative AI features
  2. Google Search Central: JavaScript SEO basics
  3. Google Search Central: Search Essentials
  4. Google Search Central: General structured data guidelines

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team