Short answer. AI search visibility depends on the same technical foundation as ordinary search: crawlable URLs, main content present in the served HTML, correct status codes, self-consistent canonicals, real internal links, an accurate sitemap, and pages that render without a browser. Generative features run on the same crawling and indexing systems as ranked search, so a page an engine cannot fetch and parse is invisible to both — and this failure is undetectable from inside a browser, where everything looks fine.
Key takeaways
- Retrieval is a hard filter: a page that is blocked, orphaned, noindexed, or misrendered never enters the candidate pool an AI answer can cite from.
- Checks have a strict order — robots rules, status codes, canonicals, discoverability — because a failure at an early step makes later content diagnostics meaningless.
- Heavy client-side rendering is the single most common cause of an otherwise good page being invisible to non-browser fetchers.
- Bot-management and CDN defaults can silently serve AI agents a shorter page or a block page, a failure only a byte-count comparison across user agents reveals.
- Page speed and Core Web Vitals are not a citation factor on their own, but the rendering work behind slow pages often causes the same content to be missing from the served HTML.
Why does crawlability come before anything else?
Retrieval acts as a filter with a hard edge, and nothing downstream of it matters until a page passes through. Crawlability is a page's basic reachability to an automated fetcher — whether it is allowed, discoverable, and returns usable content and status codes. A page that is blocked, orphaned, noindexed, returning the wrong status code, or hidden behind a rendering path a fetcher cannot follow never enters the candidate pool. Evidence density, passage structure and schema have no effect on a page that was never a candidate in the first place.
The order of checks matters, because a failure at an early step makes later diagnostics meaningless:
robots.txt— is the path allowed, for every agent that matters, not only for a wildcard?- Robots meta and
X-Robots-Tag— is the pagenoindexby accident? Staging directives shipped to production are a routine cause. - Status codes — 200 for content, single-hop 301 for moves, and no soft 404s returning 200 with an error page.
- Canonical — self-referential on the canonical URL, and consistent with what the sitemap and internal links point at.
- Discoverability — is the page linked from another crawlable page, or reachable only through a search box or a JavaScript filter?
Only after all five pass is a content diagnosis worth running. Many of the same defects covered in this technical AEO checklist show up first as a crawlability failure rather than a content gap.
What should JavaScript sites verify?
The served HTML — the markup a fetcher receives before any client-side script runs — is what decides whether content exists for a non-browser agent, and checking it in a browser hides the problem because a browser executes the JavaScript that a fetcher may not.
The test is simple: load the page with JavaScript disabled, or fetch the raw HTML directly. What remains is approximately what a fetcher that does not execute JavaScript receives. In that raw HTML, look for the main content in full rather than a loading shell or placeholder; real anchors with href attributes, since navigation built only from click handlers hides the link graph that carries authority to deep pages; title, canonical and meta directives, which can arrive missing or wrong when injected at runtime; numbers that are actually numbers rather than counters that animate up from zero and render as "0" in the served HTML; and content behind tabs and accordions, where collapsed is fine but conditionally rendered is not, because a component that mounts a panel's contents only when opened serves a fetcher a different answer than it serves a human.
Where rendering is required to see the content, server-side rendering or pre-rendering is the fix. It is not a visibility guarantee — relevance, evidence and competition still decide the outcome — but it removes a class of failure no content work can compensate for.
Are the AI agents being served the same page you are?
No page check is complete until it accounts for the possibility that a specific agent receives less than a browser does. A page can pass every crawlability and rendering check for a browser and still arrive truncated at a given AI user agent, because rate limiters, bot-management rules and CDN defaults commonly return a shorter page, a challenge page, or a 403 to agents they do not recognise — and none of that is visible from inside the site.
The verification is a byte-count comparison: bytes served to the agent in question, divided by bytes served to a browser fetch of the same URL. Run it per relevant user agent against a handful of important pages; anything materially below 1.0 is a delivery defect, not a content one. Two points are worth flagging separately. Training crawlers and retrieval crawlers are different bots doing different jobs, and common defaults block them under one toggle — blocking a retrieval fetcher makes citation impossible, a distinction covered in more depth in this guide to AI crawlers and AI search visibility. And a managed bot rule a team did not configure is still that team's rule: platform and CDN defaults change without a deploy on the publisher's side, so this check belongs on a schedule rather than a one-time launch item.
How do internal links support AI-search discovery?
Internal links make deep pages reachable and tell a crawler, through anchor text, what the destination page actually resolves. A pillar page linking out to focused pages, each linking back, puts the topic structure inside the site itself rather than only in an editorial plan.
| Pattern | Effect |
|---|---|
| Pillar page linking to focused pages, each linking back | Topic structure exists in the site, not only in the editorial plan |
| Anchor text stating the question the destination answers | The link describes the destination's purpose |
| Pages reachable only via search or a filtered list | Effectively orphaned for anything that does not execute the filter |
| Links generated client-side | Absent from the served HTML |
| Deep pages more than three clicks from an entry point | Crawled less, refreshed less |
A cluster that exists only as a content calendar is not a cluster; if the pages do not link to each other in the markup, the relationship is invisible to everything except the person who planned it. This same structural gap is what shows up in structured data and entity identity work when markup describes relationships the HTML itself never expresses.
Do page speed and Core Web Vitals matter here?
Page speed matters for users and for ordinary search quality, but it is not a citation switch on its own. There is no evidence that a faster page is more likely to be quoted by an AI answer, so treating performance work as AEO work misallocates effort.
The version of this that does matter is the shared cause behind both problems: heavy client-side rendering slows a page and removes content from the served HTML at the same time, and layout that injects the answer late hurts both outcomes. Optimising images, scripts and fonts is worth doing on its own merits, and stripping useful content to hit a performance score is not. The one performance-adjacent item with a direct effect on citability is crawl efficiency — a site that times out under a fetcher's rate limit gets less of itself fetched, regardless of how fast it loads for a human visitor.
What should a technical AI-search audit include?
A technical AI-search audit is a template-level review that checks whether a site's delivery mechanics let content be fetched, parsed and indexed at all. Run these checks as a set, on templates rather than on individual pages, because one template defect affects every page built from it:
- Index coverage and the reasons for exclusion
- Served-HTML completeness on each page template
- Byte-count parity across relevant user agents
- Canonical consistency between page, sitemap and internal links
- Redirect chains reduced to single hops
- Status codes, including soft 404s
- Internal-link depth and orphan detection
- Sitemap accuracy — canonical, indexable URLs only, with real
lastmodvalues - Structured data that matches the visible page and does not contradict it
- Titles and descriptions present in the served HTML
- Publication and modification dates that reflect actual changes
Fixing template-level defects before rewriting articles is the more efficient order: one broken template is cheaper to repair than forty rewritten pages, and it is more often the actual cause of a citation gap than the writing described in what actually gets you cited by AI answer engines.
How does Lifewood sequence this work?
Lifewood treats delivery as a gate that runs before content work, on client sites and its own, rather than as one item on a longer list. The sequence is fixed: entity resolution, then machine readability, then publishing, and pages are checked as served rather than as rendered — JavaScript disabled, byte counts compared per agent, tab and accordion contents confirmed present in the markup — because this is the failure that silently invalidates every content decision made downstream of it.
The reason the order is non-negotiable is measurement. Publishing into an unresolved delivery defect produces no movement and no way to tell whether the content itself was the problem. Establishing a baseline, fixing delivery, publishing, and re-measuring is what keeps a result attributable to a specific change rather than lost in an unmeasured backlog of defects, an approach detailed further in Lifewood's AEO and GEO service overviews.