Skip to main content
AEO/GEO

The First 90 Days of an AI Visibility Programme

August 2026 · 11 min read · Updated September 2026

Short answer. Most AI visibility programmes start by publishing, which is the wrong end. The first thirty days should establish whether the engines can reach you at all and what they currently say. The second thirty should build a baseline that survives roughly 79% day-to-day source churn on ChatGPT. Only the third should change anything, and only where an untouched control group can prove it. Ninety days is enough to know where you stand, not to change it.

Key takeaways

  • The first thirty days of an AI visibility programme verify crawler access in server logs, resolve entity identity, and freeze a set of 30–50 buyer questions with labelled archetypes.
  • Days 31–60 build a per-engine, per-market baseline with a stated sample size, because GetMentions measured day-over-day source churn of 88.3% on Gemini, 79.2% on ChatGPT and 44.4% on Perplexity.
  • Days 61–90 refresh existing pages, add sourced specificity and start third-party work, while an untouched control group of comparable pages is held back.
  • Muck Rack found 84% of AI citations across ChatGPT, Claude and Gemini were earned media, so a plan made only of your own pages has scoped itself to the minority of the outcome.
  • A day-90 report that shows a clear improvement is more likely measuring noise than success; the credible headline is the baseline, the noise floor, what changed, and what the control did.

Why does the usual first-quarter plan fail?

The standard first quarter of an AI visibility programme is a content calendar, and it fails for three reasons that are all detectable before a single page is written. Each is cheaper to check than to fix afterwards.

An AI visibility programme is a measured, sequenced effort to establish whether AI answer engines can reach a brand, what they say about it, and whether changes to owned and third-party content move those answers against a control. The sequence below avoids all three.

  • The site may not be reachable. Since 1 July 2025, new domains on Cloudflare have blocked known AI crawlers, including GPTBot, ClaudeBot and PerplexityBot, by default under a single "Block AI Bots" setting that did not distinguish training crawlers from retrieval crawlers. Cloudflare announced in July 2026 that from 15 September 2026 it splits bots into Search, Agent and Training categories, with Search allowed by default, which changes the default but not the need to check it. Content published behind a block is invisible for a reason nobody in the marketing team chose.
  • There is no baseline to move. With roughly 79% of ChatGPT's cited sources changing overnight, a programme with no pre-measurement can never demonstrate a change against noise. It can only assert one.
  • The larger half of the problem is not on your site. Muck Rack's May 2026 analysis of more than 25 million links cited by ChatGPT, Claude and Gemini found 84% were earned media, a share that has ranged from 82% to 89% across three editions since July 2025. A plan consisting entirely of your own pages has scoped itself to the minority of the outcome, which is why third-party brand mentions matter for GEO from the first week.

Days 1–30: can the engines reach you, and what do they say?

The first month establishes three things in order: whether retrieval crawlers can fetch the site, whether a machine can resolve who the brand is, and which questions the programme will be measured on. None of them involves publishing anything.

Week 1: access

Access is free to check, and it is frequently the entire problem. The full list of which AI crawlers to allow and which to block is longer than most teams expect, but the check itself has four steps.

  1. Fetch /robots.txt and check each AI user agent by name, not by wildcard. OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot are the ones that can produce a citation.
  2. Check the CDN, WAF and bot-management rules separately. A permissive robots.txt is meaningless if the edge returns 403.
  3. Confirm real hits in server logs. Permission is a claim; a 200 in the log is evidence.
  4. Check whether the text you want quoted survives with JavaScript disabled.

Week 2: identity

Establish whether a machine can resolve who you are: one canonical entity home, Organization schema with a stable @id, sameAs links to external records, consistent naming across every property, and named authors with checkable identities. An engine that cannot resolve your entity will reach for one it can.

Weeks 3–4: the question set

This is the decision that determines every number you report for the next year.

  1. Write 30–50 questions a buyer would actually ask an assistant, not keywords.
  2. Cover the archetypes deliberately and label them: recommendation and shortlist, comparison and alternatives, research and how-to, plus accuracy questions.
  3. Fix the proportions and freeze the wording.
  4. Add control questions in an adjacent category you do not intend to win.
  5. Write them natively per market. Translating an English list measures your translation.

A question registry is the frozen, archetype-labelled set of buyer questions an AI visibility programme is measured on, versioned whenever the wording changes. Archetype mix is not a detail. Analyze, across 22,295 AI answers from 460 B2B prompts and 37 organisations, found mention rate varied from 41.2% for recommendation prompts on Perplexity to 24.5% for research prompts on ChatGPT, a spread of 16.7 percentage points. Peec AI's study of 37,804 responses from 1,754 prompts (Landwehr, Ehrlinspiel and Rudzki, SSRN, June 2026) found visibility stayed stable while prompts remained above roughly 0.50 to 0.60 cosine similarity to the original, depending on the engine, and that prompts drifting into the lowest similarity band lost about half their observed visibility, while prompt length had effectively no effect.

Because the mix moves the reported number by more than sixteen points, the set has to be fixed before the first run and left alone. Choosing it after seeing early results is how a programme reports its own selection bias as progress.

Days 31–60: how do you build a baseline that survives the noise?

Run the frozen question set repeatedly, per engine, with retrieval on and off, and record five things per run rather than one. Twice weekly across four weeks on a 40-question set gives a few hundred observations per engine, enough to separate a step-change from ordinary variance.

The noise floor is the day-to-day change in cited sources that an AI visibility programme would observe even if nobody touched anything, and it is what a baseline has to be measured against. It is worth stating in the plan so nobody is surprised by it later. GetMentions, measuring 530,875 citations across 2,398 queries over seven consecutive days in June 2026, found day-over-day source churn of 88.3% on Gemini, 79.2% on ChatGPT, 75.9% on Google AI Mode and 44.4% on Perplexity, with sources cited on all seven days at 0.4%, 1.1%, 2.6% and 11.1% respectively. Sampling cost is therefore not equal across engines, and a guide to measuring AI visibility without fooling yourself starts with that fact.

Record five things per run, not one: whether the brand was named, whether a URL of yours was cited, which other brands appeared, which domains were cited, and whether what was said was accurate. A mention-rate dashboard scores a confidently wrong sentence about your pricing as a success.

Map the citation graph for your category. The domains cited in answers about your category are the actual competitive set, and they are usually not your competitors. The AI Platform Citation Source Index 2026, a synthesis of six citation studies covering more than 680 million citations, puts the top 15 domains at roughly 68% of all citations produced across ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews.

Establish which game you are in. Semrush, with Kevin Indig, tracked 1,094 US categories in ChatGPT between January and June 2026: 15.2% had a clear owner, 31.2% an emerging leader and 53.7% were unsettled, with clear owners holding the top spot in 90.4% of month-over-month comparisons. If one brand appears in more than about half of your runs, the category has an owner and displacement is slow. If the field is wide and the leader sits under a third, it is unsettled, which most categories are, and who owns a category in AI answers is a question the baseline should answer explicitly.

Days 61–90: what should you change, and how do you prove it?

Only now is there a baseline to move against, and the three workstreams should run in order of evidence quality with an untouched control group held back throughout. Refresh existing pages first, add sourced specificity second, and start the slow third-party work last.

  1. Refresh before publishing. Re-verify every figure on the pages that already answer buyer questions, move the direct answer into the first 30% of the page, phrase headings as the questions themselves, and add coverage of adjacent sub-questions. Running this as a repeatable process is what content refresh operations for AI search describes.
  2. Add sourced specificity everywhere. This is the only intervention in the category with a controlled result behind it. Aggarwal et al., "GEO: Generative Engine Optimization" (ACM SIGKDD 2024), benchmarked content changes across a 10,000-query benchmark and found that Quotation Addition raised visibility by roughly 40% and Statistics Addition by roughly 30%, with Cite Sources in the same band. These outperformed rewriting, simplification and keyword work, and keyword stuffing offered little to no improvement over making no change at all.
  3. Start the third-party workstream. Audit what the cited domains currently say about you, and separate wrong, outdated and missing into different fixes. It is slow, it addresses the majority of citations, and it will not show a result inside ninety days.

Two figures shape where the effort goes. AirOps' 2026 State of AI Search report found that for commercial queries, about 83% of AI citations came from pages updated within the previous twelve months and more than 60% from pages refreshed within six. Kevin Indig's analysis of 1.2 million ChatGPT answers and 18,012 verified citations found 44.2% of citations came from the first 30% of the cited page.

Hold an untouched control group of comparable pages. Without one, any movement is indistinguishable from the models changing underneath the benchmark, and the programme's first report is an assertion with a chart attached.

What should the day-90 report say?

The day-90 report states where the brand stands per engine and per market, what the noise floor is, what was changed, and how the untouched control group moved over the same period. Every rate in it carries a sample size.

Section Content What makes it credible
Access Which crawlers can reach the site, verified in logs Log evidence, not robots.txt alone
Baseline Mention and citation rate per engine, per market, with sample size An N next to every rate
Accuracy How often what was said about you was correct Scored on answer text, not on mentions
Category structure Owned or unsettled, with the leader's concentration The distribution of all brands named
Citation graph The domains actually cited about your category A full ranked list
Work done What changed, on which pages, when An audit trail with named reviewers
Control How the untouched set moved over the same period Reported alongside, never omitted

A day-90 report showing a clear improvement is more likely to be measuring noise than success. Ninety days is a baselining exercise, and the honest headline is: here is where we stand, here is the noise floor, here is what we changed, and here is what the control did.

What can ninety days not do?

Ninety days cannot change what a model says with search turned off, cannot show a result from third-party work, and cannot buy a guarantee from any engine.

  • It does not move memory mode. Answers with search off change when a model is retrained, on the provider's schedule.
  • Third-party work does not report inside a quarter. It is the majority of the citations and the slowest workstream.
  • No engine offers a guarantee. There is no submission, no index request and no paid placement.
  • The baseline itself will drift. Models change beneath the benchmark, which is exactly what the control questions exist to detect.

How does Lifewood run the first ninety days?

Lifewood runs the access check before quoting for content, because a site that retrieval crawlers cannot fetch cannot be improved by anything written for it, and the check costs nothing. Where the block is at the edge rather than in robots.txt, that is usually the finding.

The question registry is authored per market rather than translated, frozen at the start of a series, and versioned when it changes. Runs are reported per engine and per market with retrieval and memory kept apart, raw answers retained, and control questions run alongside the real set. 100+ languages and 40+ delivery centres across 30+ countries are what make native authorship of the registry a staffing decision rather than a translation line.

Day-90 reports are written to show the control group next to the treated pages, including when the two moved together. The measurement side of this work is delivered through Lifewood's answer engine optimisation service, and the content and third-party workstreams through its generative engine optimisation service. Buyers comparing providers on whether they hold a control group at all can start with the best AEO and GEO agencies compared.

Frequently asked questions

Days 1–30: verify crawler access in server logs, resolve entity identity, and design a frozen question set of 30–50 buyer questions with labelled archetypes. Days 31–60: run it repeatedly to build a baseline with a stated sample size, and map the citation graph for your category. Days 61–90: refresh existing content and start third-party work, holding an untouched control group.

Whether retrieval crawlers can reach your site. Since 1 July 2025, new Cloudflare domains have blocked GPTBot, ClaudeBot and PerplexityBot by default, and Cloudflare's new Search, Agent and Training categories apply from 15 September 2026. It costs nothing to check, it has to be verified in logs rather than in `robots.txt`, and it is frequently the entire problem.

Retrieval surfaces can reflect published changes in days to weeks, but proving it against roughly 79% day-to-day source churn on ChatGPT takes several weeks of repeated measurement. Third-party presence, which Muck Rack measured at 84% of AI citations across ChatGPT, Claude and Gemini, takes considerably longer than a quarter.

Thirty to fifty buyer questions, covering recommendation, comparison, research and accuracy archetypes in fixed proportions, plus control questions in an adjacent category you do not intend to win. Freeze the wording before the first run, because Peec AI found visibility stable only while prompts stayed above roughly 0.50 to 0.60 cosine similarity, with the lowest band losing about half.

By being reachable to retrieval crawlers, resolvable as an entity, and specific on the page. The GEO paper by Aggarwal et al. found Quotation Addition raised visibility by roughly 40% and Statistics Addition by roughly 30%, while keyword stuffing offered little to no improvement. Most citations are earned on third-party domains, which takes longer than a quarter.

Refresh before publishing. Recency is heavily rewarded: AirOps found about 83% of commercial-query citations came from pages updated within twelve months and more than 60% within six. The edits with controlled evidence behind them, adding sources, quotations and statistics, are cheaper to make on pages that already exist and already answer buyer questions.

Sources and further reading

  1. Cloudflare press release, 1 July 2025: Cloudflare just changed how AI crawlers scrape the internet-at-large — default block for new domains
  2. Cloudflare blog, July 2026: Your site, your rules, new AI traffic options for all customers — Search, Agent and Training categories from 15 September 2026
  3. Muck Rack, May 2026: Earned media still drives 84% of AI citations
  4. Analyze: Prompt type changes AI mention rate 17 points — 22,295 answers across 460 B2B prompts
  5. Search Engine Journal: Does prompt variance impact brand mentions? (Peec AI study) — 37,804 responses, 1,754 prompts
  6. GetMentions: AI citation volatility, a 530,875-citation study
  7. 5W: AI Platform Citation Source Index 2026 — top 15 domains at 68% of citations
  8. Semrush with Kevin Indig: AI visibility is a topic-level game — 1,094 US categories, January to June 2026
  9. Aggarwal et al., GEO: Generative Engine Optimization (arXiv:2311.09735) — ACM SIGKDD 2024
  10. AirOps: The 2026 State of AI Search — freshness of cited pages
  11. Search Engine Land: 44% of ChatGPT citations come from the first third of content — Kevin Indig's analysis, February 2026

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team