Skip to main content
AEO/GEO

Structured Data and Entity Identity: What Is Proven

August 2026 · 10 min read · Updated September 2026

Short answer. Cited pages carry JSON-LD schema roughly three times more often than uncited pages — but when Ahrefs added schema to 1,885 pages against 4,000 matched controls, the change in citations was statistically indistinguishable from zero on all three engines tested. Both results are correct. Schema is infrastructure: it removes parsing ambiguity, anchors entity identity and qualifies you for features that have documented rules. The genuinely under-invested layer is entity identity — whether a machine can resolve who you are at all.

Structured data is the most over-sold item in the AEO toolkit. It is sold as a citation lever, which the controlled evidence does not support, and dismissed as pointless, which the mechanism does not support either. This piece puts the correlation and the controlled test side by side, reconciles them, and separates what has evidence behind it from what is repeated because everyone repeats it.

Key takeaways

  • Ahrefs' controlled test added JSON-LD schema to 1,885 pages against 4,000 matched control pages and measured no statistically significant change in AI citations across Google AI Overviews, Google AI Mode, and ChatGPT.
  • Pages already cited by AI engines carry schema roughly three times more often than uncited pages, but that correlation reflects the kind of site that implements schema carefully, not the schema itself.
  • Entity identity — whether a machine can resolve an organisation to one unambiguous entity — has no published controlled effect size, but the mechanism behind it is well understood.
  • Wikipedia ranks second only to Reddit as a cited source across generative engines, per a 2026 synthesis of six studies covering more than 680 million citations, making external entity references worth maintaining.
  • Sourced, specific content placed early on a page has a controlled evidence base for AI citation lift; schema markup coverage does not.

What do the two studies actually show?

A correlation and a controlled test both exist for schema markup, and they disagree in the way correlations and controlled tests usually do.

Structured data (schema) is machine-readable markup, typically JSON-LD, that states facts on a page — a price, a date, an author, an organisation — as typed values rather than leaving them to be inferred from prose.

Study Design Result
Ahrefs controlled schema test, reported by ROI and Shine JSON-LD added to 1,885 pages, measured against 4,000 matched control pages Google AI Overviews −4.6%, Google AI Mode +2.4%, ChatGPT +2.2% — all statistically indistinguishable from zero
Ahrefs correlational scan Six million URLs AI-cited pages carried JSON-LD roughly three times more often than uncited pages

The reconciliation is not complicated. Pages that carry schema are, on average, pages maintained by organisations that also do everything else properly: technically sound, updated, structured, written by someone who cares. Schema is a marker of that population, not the cause of the citation. Schema tells you a page was built carefully. It does not make a page worth quoting.

What is schema actually for?

Discarding schema entirely is the opposite error. Structured data does three jobs, none of which is lifting citations.

  1. It removes parsing ambiguity. A price, a date, a review count or an author is stated as a typed value rather than inferred from prose, and inference is where machines get things wrong.
  2. It anchors identity. Organization, Person and sameAs are how a machine establishes that the entity on this page is the same entity as one in an external reference. That is a different problem from ranking, and content quality does not solve it.
  3. It qualifies you for features that have rules. Rich results, knowledge panels and similar surfaces have documented requirements. Those are deterministic. AI citation is not — Google's own structured data guidance is explicit that correct markup does not guarantee a rich result, and its generative AI guidance states that no special AI-only markup is required to appear in generative features.

Treat it as low-cost infrastructure with a real but narrow benefit and the decision becomes easy. For the broader technical checklist this fits into, see how to structure your website so AI engines cite you.

Why does entity identity matter more than schema coverage?

If schema is over-sold, entity resolution is under-sold, because its failure looks like a content problem rather than an identity problem.

Entity identity is whether a machine can resolve an organisation to one unambiguous entity — a single, stable record — rather than several candidates with the same or similar names.

An answer engine cannot cite a brand it cannot resolve. When a name could refer to three companies, a product line and a town, the safe behaviour for a retrieval system is to reach for a source it can resolve instead. No amount of publishing fixes that, because publishing more from an ambiguous entity adds noise to the ambiguity. We cover the surrounding entity work in more depth in entity SEO for AI search.

The mechanics are unglamorous and mostly one-time.

Signal What it does Cost
One canonical entity home Gives the entity a single authoritative URL to resolve to Low
Organization schema with a stable @id Lets every page refer to one entity rather than many Low
sameAs to external identifiers Connects your entity to records the engines already resolve Low
A Wikidata item with a stable identifier Supplies a machine-readable identifier the graph already uses Low, editorially governed
Consistent naming everywhere Prevents one organisation reading as several Ongoing discipline
Named authors with real, checkable identities Attaches expertise to a resolvable person, not a byline Ongoing

Consistency is the part that decays. Keep the official name, service descriptions, locations, leadership details and headline claims aligned across the site and the important external profiles; when facts conflict, the systems trying to connect them guess worse. Named authorship works the same way — a byline connecting to a real bio, without inflated credentials, and where no individual can be named, transparency about the editorial owner rather than an invented persona.

Wikidata and Wikipedia are community-governed with notability standards and conflict-of-interest rules. Creating or editing entries about your own organisation without disclosure breaks those rules and is routinely reverted. The legitimate route is ensuring the external references those communities require actually exist. That matters commercially because of where Wikipedia sits in the citation graph: the AI Platform Citation Source Index 2026, a synthesis of six studies covering more than 680 million citations, puts Reddit first at roughly 40% of aggregate multi-engine citation frequency, with Wikipedia second. We look at that citation graph directly in does Wikipedia still decide your AI visibility.

What does the evidence say actually moves citations?

For contrast, here is the intervention with a controlled result behind it.

Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande's GEO: Generative Engine Optimization, presented at ACM SIGKDD 2024, benchmarked content modifications across roughly 10,000 queries and nine datasets. Quotation Addition produced the largest uplift, roughly 40%, and Statistics Addition produced roughly 30%; authority-style edits — adding citations, statistics and quotations — outperformed cosmetic edits such as rewriting, simplification and keyword work, and keyword stuffing performed worse than making no change at all. Effect size varied by domain. We cover that study in detail in what gets you cited by AI answer engines.

Position on the page matters alongside it: Omnibound's 2026 AEO statistics compilation, citing SparkToro research, found 44.2% of citations came from the first 30% of a page's content.

Adding a sourced statistic to a paragraph has an evidence base. Adding FAQPage markup to that same paragraph does not. Both take about the same amount of time, and the allocation follows from that. Strong evidence here means primary research, official documentation, first-party measurement, transparent methodology and clearly sourced statistics — placed next to the claim they support, not collected in a footer. It includes limitations: a page that explains where a method fails is more credible than one claiming universal success.

Which structured-data practices have no supporting evidence?

Several practices recur in AEO proposals despite having no supporting evidence behind them.

  • Adding FAQ schema to lift citations. Correlational at best; the controlled test found nothing. Use it only where the visible page genuinely contains the questions and the site meets current eligibility rules.
  • Publishing llms.txt as a visibility tactic. Google's documentation states that Search does not use llms.txt or similar AI text files, no major crawler claims to read it, and adoption studies reported by Digital Applied found 97% of published files received zero traffic across 137,000 sites. Harmless to publish, misleading to count — and it routinely displaces the crawler-access check, which genuinely does decide whether you can be cited. We go through that check in AI crawlers: which bots to allow, which to block, and why, and cover the llms.txt question on its own in what is llms.txt, and does your website need one.
  • Marking up every entity type available. Schema you cannot maintain is schema that will eventually contradict the page.
  • Schema that does not match the visible page. A mismatched FAQPage block is ignored at best and a quality problem at worst. Unsupported review ratings and fabricated author credentials are policy exposure, not optimisation.
  • Treating markup coverage as a KPI. It measures effort, not outcome.
  • Manufactured third-party mentions. Bought reviews, placed listicles and citation farms create short-term mentions and weaken the trust they were meant to build.

What does a defensible structured-data policy look like?

A defensible policy spends most of its effort on the identity layer and treats markup as hygiene rather than a growth lever.

  1. Implement the identity layer once and maintain it: Organization, a stable @id, sameAs, consistent naming, named authors. This is the part with a real mechanism behind it.
  2. Implement Article, BreadcrumbList and FAQPage only where they match the page exactly. Cheap hygiene, no claimed lift.
  3. Generate markup from the same source of truth that renders the page, so titles, authors, dates and images cannot drift apart, and validate on every publish.
  4. Stop reporting markup coverage as an AI visibility metric. Report citation rates instead.
  5. Spend the freed effort on sourced specificity, front-loaded. That is the intervention with a controlled result behind it.
  6. Monitor whether AI systems describe you accurately. A wrong description signals that the web evidence is incomplete, inconsistent, or outweighed by a stronger third-party source.

Three honest limits apply. One controlled study is one controlled study — the Ahrefs test measured adding schema to pages that were already visible, and does not rule out an effect on pages with no other signals. Requirements for deterministic features are real and separate, so losing a rich result because markup was removed is a genuine cost unrelated to AI citation. And entity resolution has no published effect size: the mechanism is well understood, a controlled quantification is not, and this article does not claim one.

How does Lifewood put structured data and entity identity into practice?

Lifewood implements the identity layer as a one-time engineering task and then treats markup as hygiene rather than a line item with a projected citation lift attached.

Schema is generated from the same source of truth that renders the page and checked against the visible text on every publish, because markup that contradicts the copy is worse than none. The effort freed by not chasing markup coverage goes into sourced specificity and into the entity work that has a mechanism: one canonical home, consistent naming, external references that resolve. Across 100+ languages and 40+ delivery centres across 30+ countries, naming consistency is the hardest part to hold, because an organisation described three ways in three markets reads as three entities. Lifewood's own AEO services and GEO services pages apply this same identity-first approach before any markup is added.

Frequently asked questions

Not causally, on the evidence available. Ahrefs added JSON-LD to 1,885 pages against 4,000 matched controls and measured −4.6%, +2.4% and +2.2% across three engines, all statistically indistinguishable from zero. Cited pages do carry schema about three times more often, but that reflects the kind of site that implements schema rather than the schema itself.

Yes, as infrastructure. It removes parsing ambiguity, anchors entity identity, and qualifies pages for deterministic features such as rich results. What it should not be is a line in an AI visibility plan with a projected citation lift attached.

It is whether a machine can resolve your organisation to one unambiguous entity rather than several candidates. An engine that cannot resolve who you are will reach for a source it can resolve instead, and publishing more content from an ambiguous entity adds noise rather than clarity.

Wikipedia is the second most-cited domain across generative engines on the 2026 Citation Source Index synthesis of six studies and more than 680 million citations, so being described there accurately shapes many answers. Both platforms are community-governed with notability and conflict-of-interest rules, so the legitimate route is ensuring the external references they require exist, not editing entries about yourself.

Only as hygiene, and only where it matches the visible page exactly and the page meets current eligibility requirements. The specific claim that FAQ markup lifts AI citations is not supported by the controlled evidence, and mismatched markup is worse than none.

No. Google's documentation states that its Search systems do not use llms.txt or other special AI text files, and adoption studies reported by Digital Applied found 97% of published files received zero traffic across 137,000 sites. Publish it if you want; do not count it, and do not let it displace the crawler-access check.

Replacing vague claims with sourced, specific ones, near the top of the page. The ACM SIGKDD 2024 GEO study found that adding citations, statistics and quotations outperformed rewriting, simplification and keyword work, with keyword stuffing performing worse than making no change at all, and Omnibound's compilation found 44.2% of citations came from the first 30% of a page's content. Nothing guarantees a citation.

Sources and further reading

  1. Ahrefs — We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved — the controlled test and correlational scan.
  2. Search Engine Journal — Schema Markup Didn't Move AI Citations In Ahrefs Test
  3. ROI and Shine — Schema Markup for LLM Citation: Infrastructure, Not a Growth Hack
  4. Aggarwal et al., GEO: Generative Engine Optimization (arXiv:2311.09735) — ACM SIGKDD 2024.
  5. Omnibound — Answer Engine Optimization (AEO) Statistics 2026 — position-in-page citation figures.
  6. 5WPR — AI Platform Citation Source Index 2026 — synthesis of six studies covering more than 680 million citations.
  7. Google Search Central — structured data general guidelines
  8. Google Search Central — generative AI and Search guidance

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team