Concept · The Agentic Shelf

The decision-maker has moved.

For twenty-five years, commerce measurement rested on one assumption: that the human makes the final choice. In multi-turn AI, the model does — on a shelf brands can’t see. A structural framework, grounded in independent research.

By Tim de Rosen & Paul Sheals AIVO Standard · WP-2026-20 Published July 2026 ≈ 7 min read

Every commerce measurement discipline of the past twenty-five years — SEO, conversion optimisation, and lately generative engine optimisation — rests on an assumption so obvious it was never defended: that the consumer performs the terminal act of comparison and choice. Search returned ranked options; the human read, compared, and chose.

Conversational AI has moved that terminal act. When a shopper asks an assistant what to buy, the model performs the retrieval, comparison and narrowing internally, across a conversation — and hands back a recommendation already formed. The consumer’s residual role is to accept it or decline it. Measuring a brand’s position in a chain the human no longer controls at its endpoint measures the wrong thing.

The framework

The third shelf

The Agentic Shelf is the third in a sequence of commercial discovery environments. The first two share a property the third breaks: the brand can see where it stands.

Then

The physical shelf

The store aisle. The shopper compares products in front of them and picks one.

Brand can see & buy its placement
For 25 years

The digital shelf

The search results page. The shopper scans ranked options and clicks through.

Brand can measure its rank
Now

The agentic shelf

The AI conversation. The model retrieves, compares and narrows internally, across turns the brand never sees.

Brand can’t observe it at all
What to measure now

Three questions, not one

In the search era, three properties were lumped together as “visibility” because they were rarely distinct. On the Agentic Shelf they come apart — and today’s tools answer only the middle one.

Possession

Does the model know the brand exists?

Whether your brand is in the model’s knowledge at all.

Measured by: direct knowledge probes
Layer 2 · Mention

Does the model cite or name the brand?

Whether you appear in a generated response.

Measured by: GEO & today’s AI-visibility tools — this is the only axis they cover
Layer 3 · Activationthe unmeasured axis

Does the brand survive to the recommendation?

Whether you’re still standing at the moment the model actually decides — the axis that determines the commercial outcome.

Measured by: multi-turn, decision-stage probes

The empirical basis is the Linkage Gap: across 1,427 probes, brands were recognised 95.7% of the time, yet 87.3% were displaced before the final recommendation — proving that possession and mention were never the binding constraint.

Independently grounded

Why the displacement happens — and it isn’t only us

The Agentic Shelf is consistent with, and substantially explained by, four established lines of research in the AI literature — none of them conducted by us.

Liu et al., 2024

“Lost in the middle”

Models retrieve best from the start and end of context, worst from the middle — so brand facts introduced early sit exactly where retrieval is weakest by the final turn.

Laban et al., 2025 · MS + Salesforce

Multi-turn collapse

A 39% average performance drop from single-turn to multi-turn across 15 models and 200,000+ conversations. Once a model commits early, it rarely recovers.

Aggarwal et al., 2024

The limits of GEO

The paper that coined GEO shows it lifts visibility within a single response — but explicitly only at the mention level, not multi-turn survival. Necessary, not sufficient.

EMNLP 2024 · Incumbent Advantage, 2026

Incumbent bias

Models systematically favour established incumbents regardless of a competitor’s optimisation — and when everyone adopts GEO, the gains cancel out and the model reverts to favouring the incumbent.

A necessary distinction

Two kinds of displacement

Not all displacement is the same — and telling them apart decides which remedy will work. A counterfactual test (reintroducing a brand fact at the moment of recommendation) separates the two.

Addressable

The Linkage Gap

A retrieval and activation failure — the model had the fact but didn’t carry it to the decision turn. Fixable with structured evidence infrastructure that puts the right fact where the model can reach it. Most displacement falls here.

Not addressable by evidence

The Reasoning Gap

A structural judgement that a brand doesn’t fit the buyer — which persists even when the fact is available. Evidence alone won’t override it; it needs repositioning. A smaller share of cases.

Conflating the two is how a real finding gets misread as “the model is just wrong.” It isn’t — and the distinction is what tells an operator which problem they’re looking at, and which fix will move it.

The full paper

Read the complete working paper

This article is an overview. The full paper sets out the three-axis architecture, the complete related-work review, the Linkage Gap / Reasoning Gap distinction and the stated limitations — published open-access on Zenodo with a permanent DOI.

Citation: de Rosen, T. & Sheals, P. (2026). The Agentic Shelf: A Structural Framework for Brand Decision Architecture in Multi-Turn AI Recommendation Systems. AIVO Standard Working Paper WP-2026-20. Zenodo. https://doi.org/10.5281/zenodo.21131113 · Licensed CC-BY-4.0.