Skip to main content
AEO/GEO

Should You Serve Content to AI Agents Through an API or MCP Server?

October 2026 · 12 min read

Short answer. For most organisations, not yet as a replacement, and increasingly yes as an addition. Answer engines still build citations by crawling, so a readable website remains the foundation of AEO and GEO. An API, MCP server or NLWeb endpoint is worth prototyping only where your content is queryable, high-volume or transactional.

Key takeaways

  • Crawling is an inefficient way to move facts between machines, and Cloudflare's crawl-to-refer data shows AI platforms requesting thousands of pages per referred visit, although those ratios have fallen through 2026.
  • The Model Context Protocol moved to the Linux Foundation's Agentic AI Foundation in December 2025, with backing from competing platforms, which makes it a standard that can be built against without betting on one vendor.
  • An llms.txt file points to content, an API serves data, MCP standardises how a model connects to tools, and NLWeb turns existing Schema.org data into a queryable endpoint that also speaks MCP.
  • No published, controlled evidence shows that an MCP server lifts citations in public answer engines, so its value case is queryable, transactional, oversized or enterprise-consumed content.
  • Sequencing decides the return: readable HTML and accurate structured data come first, and an agent interface comes second as an addition rather than a replacement.

What problem is an agent interface actually solving?

An agent interface solves the inefficiency of crawling, where an agent wanting one fact such as a price or a stock level must fetch a whole page and infer structure that was never declared. Handing the agent the data directly is either a better deal for both sides or an unnecessary second surface to maintain, depending on what your content is.

An agent interface is a machine-readable endpoint that lets an AI agent request specific data or actions directly, instead of reconstructing them from web pages.

Crawling is a lossy way to move information between two machines. A page is a document formatted for a human eye; an agent wanting one fact — a price, a stock level, a specification, an appointment slot — has to fetch the whole document, parse it, and infer structure that was never declared.

The inefficiency shows up in the traffic data. Cloudflare publishes a crawl-to-refer ratio comparing pages an AI platform's crawlers request against visits that platform refers back, and the spread is enormous. For one sample week in June 2025, the ratios ran from roughly 70,900:1 at one end to 0.1:1 at the other. Cloudflare's own year-in-review reporting showed Anthropic's ratio falling by 87% across 2025 and still sitting at around 38,000 pages crawled per referred visit in July 2025.

Those ratios have continued to compress through 2026 as assistants matured into products that return clicks, and third-party reconstructions of the Radar data vary considerably depending on the window chosen, which is itself the point. The metric is volatile, the direction is towards parity, and the underlying asymmetry is real: crawling costs the publisher and delivers the platform far more than it returns. Which crawlers to allow in the first place is covered in which AI bots to allow and which to block.

Crawl-to-refer asymmetry, as reported by Cloudflare

Reporting window Platform Pages crawled per referral
June 2025 sample week Highest observed ratio About 70,900 : 1
July 2025 Anthropic About 38,000 : 1 (down 87% across 2025)
June 2025 sample week Lowest observed ratio 0.1 : 1

The ratio is highly volatile between windows and has fallen through 2026, so treat any single figure as a snapshot, not a constant. Source: Cloudflare Radar and Cloudflare blog, 2025.

How settled is the protocol layer in 2026?

The protocol layer is more settled than it looked a year ago. In December 2025 Anthropic donated the Model Context Protocol to the Agentic AI Foundation under the Linux Foundation, so MCP is now governed neutrally and supported by competing platforms.

The Model Context Protocol (MCP) is an open standard that defines how an AI model discovers and calls external tools, resources and prompts through a client-server pattern.

On 9 December 2025, Anthropic donated MCP to the newly formed Agentic AI Foundation, a directed fund under the Linux Foundation. The MCP project's own announcement put the protocol at over 97 million monthly SDK downloads, around 10,000 active servers, and first-class client support across ChatGPT, Claude, Cursor, Gemini, Microsoft Copilot and Visual Studio Code. The Linux Foundation announcement named MCP, Block's goose and OpenAI's AGENTS.md as founding projects, with platinum members including AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI.

That matters for a reason that has nothing to do with the technology. A standard backed by competitors and governed neutrally is a standard you can build against without betting on a vendor. A year earlier, the reasonable objection to MCP was that the agentic web had several competing standards and therefore none, a scepticism publishers voiced openly at the time (Digiday, 2025). The consolidation has partly answered that, though the picture remains layered rather than unified: MCP for backend tools, WebMCP for browser-side interfaces, A2A for agent-to-agent communication, NLWeb for content queries (analysis, February 2026).

The agent-facing stack by maturity and effort

Surface Maturity Build effort Our assessment
Readable HTML Established Low Do this first
Structured data Established Low Do this first
Public REST API Established Medium to high Build if the use case justifies it
MCP server Fairly established High Build if the use case justifies it
NLWeb endpoint Emerging Medium Prototype
llms.txt Emerging Very low Cheap but unproven
WebMCP Early draft High Watch

Positions are our assessment based on the adoption and governance evidence cited in this article, not a measured benchmark.

What is the difference between llms.txt, an API, MCP and NLWeb?

llms.txt is a pointer file, an API serves structured data, MCP standardises how a model connects to tools and data, and NLWeb turns a site's existing structured data into a natural-language query endpoint that also acts as an MCP server. They do genuinely different jobs and are often discussed together.

NLWeb is an open-source Microsoft project that turns a website's existing Schema.org data into a natural-language query surface, with every instance also acting as an MCP server.

llms.txt is a Markdown file at the root of a site listing its most important pages for AI systems. It is a community proposal rather than a ratified standard, and it is a pointer, not an interface: it tells a system what to read, and nothing happens unless the system chooses to read it. Google states plainly in its site-owner guidance that new machine-readable AI files are not required to appear in its AI features. The format is explained in what llms.txt is and does.

A public REST API is the oldest answer and still a good one. It serves structured data to any client, is well understood by engineering teams, and requires no new protocol. Its limitation is discovery and semantics: an agent needs documentation and a schema to know what your endpoints mean.

MCP standardises how a model connects to tools and data through a client-server pattern, defining tools the model can call, resources it can read, and prompts it can invoke, with built-in discovery so a client can enumerate what a server offers without bespoke integration. Write one server, and compatible clients across vendors can use it.

NLWeb sits a layer up. Its central claim is leverage: Schema.org and related formats already exist on a very large share of the web, and NLWeb uses them as a semantic layer, with every instance also acting as an MCP server exposing an ask method. Microsoft's framing is that NLWeb is to MCP roughly what HTML is to HTTP.

The practical ordering follows from that. Readable HTML and accurate structured data are prerequisites for everything else, including NLWeb, which is built on them; the groundwork is described in structured data and entity identity for AEO. An llms.txt file is cheap and unproven. An MCP server is real engineering with real maintenance, and it is worth doing when there is a genuine query surface behind it.

Does exposing an endpoint actually improve AI visibility?

No published, controlled evidence shows that standing up an MCP server increases citations in ChatGPT, Gemini or Perplexity. The major answer engines build their citation sets primarily from crawled and indexed web content, so a brand that replaced its pages with an MCP server would disappear from AI answers.

An honest article has to slow down here. How the engines choose sources is set out in how ChatGPT picks its sources, and none of it depends on an agent endpoint.

The sceptical case was put well by one publishing executive early in MCP's rise: it is not obvious that a server gives agents information they could not already obtain by crawling, and to be worth using, the interface has to offer something more valuable than scraping the site (Digiday, 2025). That test, more valuable than scraping, is the right one, and it has clear answers in some cases and none in others.

It passes when the content is:

  • Queryable rather than readable. Availability, inventory, seat counts, timetables, rates, coverage by postcode: facts that live in a database and change.
  • Too large to crawl sensibly. A catalogue, a documentation set, an archive where the useful unit is a query result rather than a page.
  • Transactional. Booking, quoting, configuring: actions rather than facts.
  • Consumed inside an enterprise. Where the users are your own customers' internal copilots rather than public answer engines. This is the quietest and strongest current use case, and it does not depend on public AI search at all.

It fails when the content is a set of marketing pages and blog posts. An MCP server wrapping content that a crawler could read in full adds maintenance without adding reach.

Who should build an agent interface now, and who should wait?

Build now if you sell something with live parameters or your documentation is a competitive asset. Wait if your pages are not yet fully machine-readable, because fixing rendering will move visibility far more than adding a protocol.

Build now if you sell something with live parameters — travel, logistics, financial products, software with usage-based pricing, anything where the honest answer to a buyer's question is "it depends on these variables". Build now, too, if your documentation is a competitive asset. Developer-facing companies have seen documentation become the most-cited surface they own, because dense, normalised, well-segmented reference material is exactly what retrieval systems prefer.

Wait if your pages are not yet fully machine-readable. Sequencing matters here and is routinely got wrong: a brand whose prices render client-side and whose specifications sit in images will gain far more from fixing rendering than from adding a protocol. The unglamorous work is the work that moves visibility, and the technical AEO checklist is the place to start.

Watch the browser-side layer. WebMCP is an early-stage draft specification, not a standard, with an incomplete security model and no discovery mechanism as of early 2026. That is a reasonable thing to follow and a poor thing to depend on.

One governance point deserves flagging before anyone ships. An agent interface is an access-control surface as well as a distribution surface. It determines what agents can retrieve, at what rate, under what identification, and with what logging. Treat it as infrastructure with a security review, not as a marketing asset.

How does this connect to Lifewood's own AEO and GEO work?

Lifewood Data Technology provides AEO and GEO services, so it has a commercial interest in this subject. Most engagements begin further back than agent interfaces, with whether the engines can read the site at all and whether facts stay consistent across languages and markets.

Declaring the interest: Lifewood provides answer engine optimisation (AEO) and generative engine optimisation (GEO) services. That earlier layer is where results move. We treat agent interfaces as the next layer up, worth prototyping for clients whose content is genuinely queryable, and worth deferring for those whose pages are not yet legible. For buyers comparing providers, best AEO and GEO agencies compared sets out the criteria.

The organisations doing the substantive work here are mostly not marketing agencies. The Agentic AI Foundation under the Linux Foundation now governs MCP. Microsoft maintains NLWeb as an open-source reference implementation. Cloudflare publishes the crawl and referral data that lets anyone reason about the economics rather than guess. Any serious agent-readiness programme draws on those directly.

What should you do to prototype an agent interface?

Fix rendering first, then pick one real customer question that depends on variables and expose it as a single well-documented tool. Instrument it, review access as a security matter, and keep your pages live.

  • Fix rendering first. Fetch your key pages with JavaScript disabled. If the facts are missing, stop here and fix that instead.
  • Find the query, not the page. Write down the ten questions customers ask that a static page answers badly because the answer depends on variables. That list is your endpoint specification.
  • Check what you already have. Most organisations with a public API are closer than they think; the gap is often documentation and schema, not capability.
  • Start with one tool. A single well-documented MCP tool that answers one real question beats a broad server nobody calls.
  • Use your structured data as the substrate. If Schema.org markup is already in place, NLWeb is a shorter path than building from scratch.
  • Instrument it. Log which agents call it, what they ask, and what they do with the response. Without that, you cannot evaluate whether it is worth maintaining.
  • Review access and rate limits as security, not marketing. Decide deliberately who may call it and how often.
  • Keep the pages. The interface is an addition. Crawling is still how citations are built.

Frequently asked questions

There is no evidence that it will on its own. Public answer engines build citations mainly from crawled and indexed content. An MCP server serves agents that choose to call it, which is a different distribution path from the search and retrieval systems that decide which brands appear in answers.

It is cheap and harmless, and some teams report parsing benefits. It is a community proposal rather than a standard, and Google states that new machine-readable AI files are not required for its AI features. Publish it if the effort is trivial, but do not treat it as a visibility strategy.

An API exposes endpoints that a developer integrates against one by one. MCP standardises how a model discovers and calls tools, so a compliant client can enumerate and use your server without a bespoke integration for each platform. Many MCP servers simply wrap an existing API.

NLWeb builds a natural-language query layer on the Schema.org data a site already publishes, and each instance also functions as an MCP server. For sites with good structured data it is a shorter route than designing an agent interface from scratch, though it is newer than a REST API.

It is a draft community specification with an incomplete security model and no discovery mechanism as of early 2026. Follow its progress, but do not build a programme or a roadmap commitment on it until it stabilises and gains broad browser support.

No. It sits on top. A site whose content is not machine-readable will not be rescued by an endpoint, because the citation layer still runs on crawled pages. Keep investing in readable pages and structured data, and add an agent interface only where the use case justifies it.

Sources and further reading

  1. Model Context Protocol blog: MCP joins the Agentic AI Foundation (9 December 2025) — the donation to the Linux Foundation, 97 million monthly SDK downloads, roughly 10,000 active servers, and client support.
  2. Linux Foundation: Formation of the Agentic AI Foundation (9 December 2025) — founding projects MCP, goose and AGENTS.md, and the platinum member list.
  3. NLWeb reference implementation, GitHub — NLWeb's design, native MCP support, the ask method, and its use of Schema.org as a semantic layer.
  4. Microsoft Tech Community: Optimise your site for agents — how NLWeb, MCP and llms.txt relate, and the HTML-to-HTTP framing.
  5. Cloudflare blog: AI search crawl-to-refer ratio on Radar — definition of the metric and the June 2025 sample range from roughly 70,900:1 to 0.1:1.
  6. Cloudflare blog: The crawl before the fall of referrals — 2025 crawl-to-refer movement by platform, including the 87% decline and July 2025 figures.
  7. Digiday: WTF is Model Context Protocol and why should publishers care? — the publisher-side sceptical case and the more-valuable-than-scraping test.
  8. Ivan Turkovic: WebMCP is coming (February 2026) — the layered stack of MCP, WebMCP, A2A and NLWeb, and WebMCP's draft status.
  9. Google Search Central: AI features and your website — confirmation that new machine-readable AI files are not required for Google's AI features.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team