LIFEWOOD
Ready100
AI search

Structured Data and Entity Identity: What Is Proven

Short answer. Cited pages carry JSON-LD schema roughly three times more often than uncited pages — but when Ahrefs added schema to 1,885 pages against 4,000 matched controls, the change…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. Cited pages carry JSON-LD schema roughly three times more often than uncited pages — but when Ahrefs added schema to 1,885 pages against 4,000 matched controls, the change in citations was statistically indistinguishable from zero on all three engines tested. Both results are correct. Schema is infrastructure: it removes parsing ambiguity, anchors entity identity and qualifies you for features that have documented rules. The genuinely under-invested layer is entity identity — whether a machine can resolve who you are at all.

Structured data is the most over-sold item in the AEO toolkit. It is sold as a citation lever, which the controlled evidence does not support, and dismissed as pointless, which the mechanism does not support either. This piece puts the correlation and the controlled test side by side, reconciles them, and separates what has evidence behind it from what is repeated because everyone repeats it.


What do the two studies actually show?

This is one of the few places in AEO where a correlation and a controlled test both exist, and they disagree in the way correlations and controlled tests usually do.

Study Design Result
Ahrefs controlled schema test, reported by ROI and Shine JSON-LD added to 1,885 pages, measured against 4,000 matched control pages Google AI Overviews −4.6%, Google AI Mode +2.4%, ChatGPT +2.2% — all statistically indistinguishable from zero
Ahrefs correlational scan Six million URLs AI-cited pages carried JSON-LD roughly three times more often than uncited pages

The reconciliation is not complicated. Pages that carry schema are, on average, pages maintained by organisations that also do everything else properly: technically sound, updated, structured, written by someone who cares. Schema is a marker of that population, not the cause of the citation.

Schema tells you a page was built carefully. It does not make a page worth quoting.


What is schema actually for?

Discarding it is the opposite error. Structured data does three jobs, none of which is "lift citations".

  1. It removes parsing ambiguity. A price, a date, a review count or an author is stated as a typed value rather than inferred from prose, and inference is where machines get things wrong.
  2. It anchors identity. Organization, Person and sameAs are how a machine establishes that the entity on this page is the same entity as one in an external reference. That is a different problem from ranking, and content quality does not solve it.
  3. It qualifies you for features that have rules. Rich results, knowledge panels and similar surfaces have documented requirements. Those are deterministic. AI citation is not — Google's own structured data guidance is explicit that correct markup does not guarantee a rich result, and its generative AI guidance states that no special AI-only markup is required to appear in generative features.

Treat it as low-cost infrastructure with a real but narrow benefit and the decision becomes easy.


Why is entity identity the layer that matters more?

If schema is over-sold, entity resolution is under-sold, because its failure looks like a content problem.

An answer engine cannot cite a brand it cannot resolve. When "Acme" could be three companies, a product line and a town, the safe behaviour for a retrieval system is to reach for a source it can resolve instead. No amount of publishing fixes that, because publishing more from an ambiguous entity adds noise to the ambiguity.

The mechanics are unglamorous and mostly one-time.

Signal What it does Cost
One canonical entity home Gives the entity a single authoritative URL to resolve to Low
Organization schema with a stable @id Lets every page refer to one entity rather than many Low
sameAs to external identifiers Connects your entity to records the engines already resolve Low
A Wikidata item with a stable identifier Supplies a machine-readable identifier the graph already uses Low, editorially governed
Consistent naming everywhere Prevents one organisation reading as several Ongoing discipline
Named authors with real, checkable identities Attaches expertise to a resolvable person, not a byline Ongoing

Consistency is the part that decays. Keep the official name, service descriptions, locations, leadership details and headline claims aligned across the site and the important external profiles; when facts conflict, the systems trying to connect them guess worse. Named authorship works the same way — a byline connecting to a real bio, without inflated credentials, and where no individual can be named, transparency about the editorial owner rather than an invented persona.

Wikidata and Wikipedia are community-governed with notability standards and conflict-of-interest rules. Creating or editing entries about your own organisation without disclosure breaks those rules and is routinely reverted. The legitimate route is ensuring the external references those communities require actually exist. That matters commercially because of where Wikipedia sits in the citation graph: the AI Platform Citation Source Index 2026, a synthesis of six studies covering more than 680 million citations, puts Reddit first at roughly 40% of aggregate multi-engine citation frequency and Wikipedia second, appearing in 26–48% of ChatGPT top-10 answers.


What does the evidence say does move citations?

For contrast, here is the intervention with a controlled result behind it. Aggarwal et al., "GEO: Generative Engine Optimization" (ACM SIGKDD 2024), benchmarked content modifications across roughly 10,000 queries and nine datasets: targeted content changes raised visibility in generative engine responses by up to 40%, authority-style edits — adding citations, statistics and quotations — outperformed cosmetic edits such as rewriting, simplification and keyword work, and keyword stuffing performed worse than making no change at all. Effect size varied by domain. We cover that study in detail in what gets you cited by AI answer engines.

Position on the page matters alongside it: Omnibound's 2026 AEO statistics compilation found 55% of sampled AI Overview citations came from the first 30% of the cited page.

Adding a sourced statistic to a paragraph has an evidence base. Adding FAQPage markup to that same paragraph does not. Both take about the same amount of time, and the allocation follows from that.

Strong evidence here means primary research, official documentation, first-party measurement, transparent methodology and clearly sourced statistics — placed next to the claim they support, not collected in a footer. It includes limitations: a page that explains where a method fails is more credible than one claiming universal success.


The cargo-cult list

Practices that recur in AEO proposals and have no supporting evidence behind them.

  • Adding FAQ schema to lift citations. Correlational at best; the controlled test found nothing. Use it only where the visible page genuinely contains the questions and the site meets current eligibility rules.
  • Publishing llms.txt as a visibility tactic. Google's documentation states that Search does not use llms.txt or similar AI text files, no major crawler claims to read it, and adoption studies reported by Digital Applied found 97% of published files received zero traffic across 137,000 sites. Harmless to publish, misleading to count — and it routinely displaces the crawler-access check, which genuinely does decide whether you can be cited.
  • Marking up every entity type available. Schema you cannot maintain is schema that will eventually contradict the page.
  • Schema that does not match the visible page. A mismatched FAQPage block is ignored at best and a quality problem at worst. Unsupported review ratings and fabricated author credentials are policy exposure, not optimisation.
  • Treating markup coverage as a KPI. It measures effort, not outcome.
  • Manufactured third-party mentions. Bought reviews, placed listicles and citation farms create short-term mentions and weaken the trust they were meant to build.

What does a defensible policy look like?

  1. Implement the identity layer once and maintain it. Organization, a stable @id, sameAs, consistent naming, named authors. This is the part with a real mechanism behind it.
  2. Implement Article, BreadcrumbList and FAQPage only where they match the page exactly. Cheap hygiene, no claimed lift.
  3. Generate markup from the same source of truth that renders the page, so titles, authors, dates and images cannot drift apart, and validate on every publish.
  4. Stop reporting markup coverage as an AI visibility metric. Report citation rates instead.
  5. Spend the freed effort on sourced specificity, front-loaded. That is the intervention with a controlled result behind it.
  6. Monitor whether AI systems describe you accurately. A wrong description signals that the web evidence is incomplete, inconsistent, or outweighed by a stronger third-party source.

Three honest limits. One controlled study is one controlled study — the Ahrefs test measured adding schema to pages that were already visible, and does not rule out an effect on pages with no other signals. Requirements for deterministic features are real and separate, so losing a rich result because markup was removed is a genuine cost unrelated to AI citation. And entity resolution has no published effect size: the mechanism is well understood, a controlled quantification is not, and this article does not claim one.


How Lifewood approaches this

Lifewood implements the identity layer as a one-time engineering task and then treats markup as hygiene rather than a line item with a projected citation lift attached. Schema is generated from the same source of truth that renders the page and checked against the visible text on every publish, because markup that contradicts the copy is worse than none.

The effort freed by not chasing markup coverage goes into sourced specificity and into the entity work that has a mechanism: one canonical home, consistent naming, external references that resolve. Across 50+ languages and 40+ delivery centres across 30+ countries, naming consistency is the hardest part to hold, because an organisation described three ways in three markets reads as three entities.

See AEO services, GEO services, AEO and GEO providers and the glossary.


Sources and further reading

  • Ahrefs controlled schema test (1,885 pages, 4,000 matched controls) and six-million-URL correlational scan, reported by ROI and Shine, 2026.
  • Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, GEO: Generative Engine Optimization, ACM SIGKDD 2024.
  • Omnibound, Answer Engine Optimization statistics 2026 — position-in-page figures.
  • AI Platform Citation Source Index 2026 — synthesis of six studies covering more than 680 million citations.
  • Google Search Central — general structured data guidelines, and guidance on optimising for generative AI features.
  • SE Ranking and Ahrefs llms.txt adoption studies, reported by Digital Applied, 2026.

Frequently asked questions

Not causally, on the evidence available. Ahrefs added JSON-LD to 1,885 pages against 4,000 matched controls and measured −4.6%, +2.4% and +2.2% across three engines, all statistically indistinguishable from zero. Cited pages do carry schema about three times more often, but that reflects the kind of site that implements schema rather than the schema itself.

Yes, as infrastructure. It removes parsing ambiguity, anchors entity identity, and qualifies pages for deterministic features such as rich results. What it should not be is a line in an AI visibility plan with a projected citation lift attached.

It is whether a machine can resolve your organisation to one unambiguous entity rather than several candidates. An engine that cannot resolve who you are will reach for a source it can resolve instead, and publishing more content from an ambiguous entity adds noise rather than clarity.

Wikipedia is the second most-cited domain across generative engines, appearing in 26–48% of ChatGPT top-10 answers on the 2026 Citation Source Index synthesis, so being described there accurately shapes many answers. Both platforms are community-governed with notability and conflict-of-interest rules, so the legitimate route is ensuring the external references they require exist, not editing entries about yourself.

Only as hygiene, and only where it matches the visible page exactly and the page meets current eligibility requirements. The specific claim that FAQ markup lifts AI citations is not supported by the controlled evidence, and mismatched markup is worse than none.

No. Google's documentation states that its Search systems do not use llms.txt or other special AI text files, and adoption studies reported by Digital Applied found 97% of published files received zero traffic across 137,000 sites. Publish it if you want; do not count it, and do not let it displace the crawler-access check.

Replacing vague claims with sourced, specific ones, near the top of the page. The ACM SIGKDD 2024 GEO study found that adding citations, statistics and quotations outperformed rewriting, simplification and keyword work, with keyword stuffing performing worse than making no change at all, and Omnibound's compilation found 55% of sampled AI Overview citations came from the first 30% of the page. Nothing guarantees a citation.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team