Concept · The Layer Mismatch

The category has solved the wrong problem.

The GEO industry is producing real, measurable gains — in the wrong evidence layer. First-prompt visibility and decision-stage recommendation are different problems at different layers, and improving the first doesn’t move the second. Here’s why — and why the category structurally can’t fix it.

By Paul Sheals & Tim de Rosen AIVO Standard · WP-2026-08 Published April 2026 ≈ 8 min read

Generative Engine Optimisation was founded on a correct insight: traditional search metrics don’t measure how a brand performs in AI answers, and that needs different tooling. The category built sophisticated ways to measure citation frequency and first-prompt visibility, and to improve them. That achievement is real.

But it has accurately solved one measurement problem while creating another. What it measures — first-prompt citation and brand visibility in AI responses — is a correct measurement of something real. It is not a measurement of the thing brands believe they’re buying: AI purchase-recommendation performance — whether, when a buyer actually asks an AI what to choose, the brand wins. And first-prompt visibility doesn’t predict that outcome. A brand can appear consistently in AI responses while recording zero recommendation wins at the decision stage.

“The GEO category has solved the first-prompt visibility problem. It has not solved the decision-stage recommendation problem. The two are different problems at different evidence layers.”
Why

Three evidence layers, not one

AI models use different evidence for different tasks within the same conversation. At the first prompt, they reach for what’s recent and frequently cited. At the decision turn, they reason from something else entirely — and the two are structurally independent.

Layer 1
Community & editorial

What GEO populates

Reddit, LinkedIn, YouTube, blogs, press, on-site copy — the content AI retrieves at the first prompt. Real and measurable, but the most volatile layer (much of what AI cites changes month to month).

Layer 2
Editorial authority

Higher weight, still not decisive

Tier-1 press, analyst reports, awards, peer-reviewed and regulatory sources. GEO acknowledges this layer — but doesn’t measure whether it satisfies the specific criteria a model applies at the decision turn.

Layer 3
Knowledge-graph anchorsdecides the sale

What GEO cannot touch

Wikidata entity definitions, Wikipedia category statements, trained-model entity representations, structured evidence architecture — what the model reasons from when it applies criteria filters at the decision turn. The most stable layer, and the most determinative of the outcome.

Gains in Layers 1 and 2 do not propagate to Layer 3. And Layer 3 is closed to the category: Wikipedia prohibits paid editing; Wikidata requires verifiable third-party citations. There is no “knowledge-graph optimisation feature” a content platform can ship, because the platforms that constitute Layer 3 reject optimisation as a concept.

Worse, not just absent

Three ways GEO content can make it worse

The mismatch isn’t only a failure to reach Layer 3. In documented cases, GEO content actively conflicts with the Layer 3 evidence — and the model resolves the conflict by lowering its confidence in the brand.

Conflict 01

Entity-definition conflict

Layer 3 classifies a brand in one category while new content positions it in another. The model trusts the authoritative anchor for its criteria check — so the brand is visible at the first prompt but absent at the decision turn. More content deepens the conflict; it doesn’t resolve it.

Conflict 02

Claim-consistency conflict

Marketing-language claims that don’t match the neutral, critical language of Layer 3 sources read as low-confidence. The model hedges the recommendation — or routes the buyer to a competitor whose claims are internally consistent.

Conflict 03

Temporal-weighting conflict

Trained entity definitions persist beyond their sources’ currency. New content (2025–26) is layered on older trained knowledge — the old takes precedence in reasoning, the new in retrieval, and the model never reconciles the two.

Illustrative cases in the paper span B2B SaaS, cloud infrastructure, prestige beauty, luxury fragrance and travel — some brands are GEO-platform clients, some aren’t; the mechanism is identical. Brand-level cases are described here as archetypes; the findings characterise AI model behaviour, not brand quality.

Structural, not negligence

Why the category can’t fix it

This isn’t a criticism of any company — it’s the predictable result of three constraints that operate on every player in the category at once.

Constraint 01

The platforms are inhospitable

Wikipedia and Wikidata reject commercial tooling by design. The signals models weight most for entity definition are controlled by third-party editorial communities, not by brands or their agencies.

Constraint 02

The metric is the wrong metric

The category measures citation frequency because it’s measurable, improvable on a short timeline and reportable in a dashboard. Decision-stage win rate isn’t — so it goes unmeasured, and therefore unaddressed.

Constraint 03

The business model favours content

Content production is recurring, scalable and billable. Knowledge-graph remediation is a one-time, low-cadence fix that doesn’t generate the same returns — so the incentive points at Layers 1–2.

What it actually takes

A different layer needs a different intervention

If the mismatch is structural, closing it requires work at the layer that decides the outcome — not more content. That means decision-stage measurement (probing the full multi-turn buying conversation, not the first prompt), classifying why a gap exists before prescribing a fix, and Layer 3 remediation: correcting knowledge-graph entity definitions, structuring evidence into machine-extractable formats, and deploying brand.context declarations — maintained continuously, since entity definitions drift.

“The layer mismatch is not a gap that more GEO content can close. It’s a gap the wrong kind of content has been trying to close — in some cases making it wider.”

GEO solved a real problem. It just isn’t the problem its clients are paying to solve. AI purchase-recommendation performance is a Layer 3 evidence-architecture problem — and the brands that understand that before their competitors do will hold a structural advantage no amount of content spend can replicate.

The full paper

Read the complete working paper

This article is an overview. The full paper sets out the three-layer model, the three conflict mechanisms, five industry case studies, the structural constraints on the category, and the methodology required to address the mismatch — published open-access on Zenodo with a permanent DOI.

Citation: Sheals, P. & de Rosen, T. (2026). The Layer Mismatch: Why GEO Visibility Gains Do Not Translate to Decision-Stage Recommendation — and Why the Category Cannot Fix It. AIVO Standard, Working Paper WP-2026-08. Zenodo. https://doi.org/10.5281/zenodo.19840293 · Licensed CC-BY-4.0.