Primary research · AI Decision Intelligence

You can win the AI conversation and still lose the sale.

A six-month study of how AI assistants actually handle purchase journeys — and why the brands that score highest on AI visibility can be the ones quietly displaced at the moment of purchase.

Research: AIVO Edge · AIVO Standard™ 195 brands · 4 AI platforms · 3 phases Published April 2026 ≈ 9 min read

There is a new stage in the buying journey, and it sits before every channel brands currently measure. Before the Google search, before the retailer page, before the review — a consumer opens an AI assistant and begins forming a view of what they should buy.

By early 2026 this is not a future risk; it is the live commercial environment. ChatGPT reached 900 million weekly users. AI search traffic converts at around 14% versus roughly 3% for traditional search — and the AI-referred consumer arrives more decided. Between 40% and 55% of consumers in top purchasing categories now use AI specifically to make buying decisions, and McKinsey projects $750 billion in US consumer spending flowing through AI-powered search by 2028.

The most valuable stage of the funnel has migrated into a conversation most brands cannot see, measure, or influence. This research set out to map what actually happens inside it.

The unit that matters

It’s a reasoning chain, not a search result

When a consumer asks an AI about a purchase, what follows isn’t a lookup — it’s a multi-turn reasoning sequence. The model affirms, compares, applies criteria, narrows the field, and eventually offers to route the consumer somewhere to buy. That last step — the commercial handoff moment — is the fulcrum of the whole journey. The brand that is primary at that moment captures the routing, the intent, and the sale that follows.

“A visibility score tells you a brand appeared somewhere. It cannot tell you where in the sequence, under what framing, next to which competitor — or whether the brand survived to the moment of purchase.”

This is why a single visibility score — the standard measure in the AI-optimisation industry — is structurally blind to the outcome that matters. It captures a moment, not a sequence.

The central discovery

Structural narrative substitution

Across twelve brands, four platforms and five identical runs each, one pattern held with almost no variance: a natural four-turn purchase conversation reliably converts an early, genuine endorsement of a brand into a recommendation for a competitor. It runs in four stages.

1

Positive anchoring

At turn one the model validates the brand you named. The endorsement is genuine — and it is precisely what sets up the displacement, by making your brand the reference point.

2

Comparison expansion

The model broadens the conversation, introducing a competitive set the consumer never asked for. That set defines the frame — and the frame decides who can win.

3

Optimisation substitution the trigger

A criterion is introduced — “which is most clinically proven?”, “best value?” — and the model reframes the whole evaluation around it. This turn is the universal displacement trigger across every brand and platform tested.

4

Purchase lock-in

The model makes a specific recommendation — usually for whichever brand won the criterion it just chose to prioritise. Not the brand the consumer started with.

Crucially, this is not hallucination. Noise does not reproduce five times across four platforms with zero variance in direction. The model is accurately synthesising a positioning hierarchy encoded in its training data — where, in this category, peer-reviewed evidence can outrank heritage, ingredient novelty can outrank brand equity, and price efficiency can outrank prestige. The problem isn’t that the model is wrong. It’s that the hierarchy it faithfully reproduces doesn’t serve the brand’s interests.

Five headline findings

What the evidence shows

Three research phases — controlled displacement tests, a directed-probe corpus of 8,500+ sequences, and naturalistic “journey probes” that let the AI lead — produced five findings that current measurement cannot detect.

0
purchase recommendations across 48 naturalistic journey runs. AI shapes preference relentlessly — but it doesn’t “say buy”. It offers to route you.
90% → 0
a brand can win 90% of directed AI conversations and receive zero purchase recommendations in real ones. The correlation between the two is effectively nil.
4
distinct, reproducible platform failure modes — independent of brand quality, tier or positioning.
Invisible
some of the highest-visibility brands were entirely absent from their own category when a consumer asked for a recommendation without naming anyone.

That last finding is the most consequential — and the one current tools are architecturally incapable of seeing. A brand can hold strong “AI visibility” yet never enter the consideration set an AI builds for a generic category question. In the study, a decades-old heritage serum was absent from every generic recommendation on all three platforms; the AI proposed newer, citation-rich challengers instead. The consumer wasn’t shown that brand and rejected it — they were handed a shortlist it was never on.

Named-brand results from this programme are held under NDA and are not reproduced here. The findings describe AI model behaviour, not brand quality.

Platform intelligence

Four models, four failure modes

AI visibility is not one phenomenon with one fix. Each platform displaces brands differently — so a strategy that works on one can have no effect on another.

ChatGPT

The channel-first loyalist

Commits to the brand you name and spends the journey helping you buy it — but never introduces a competitor and never closes. In directed tests, though, it’s the strictest clinical evaluator. Same platform, opposite behaviours.

Gemini

The educational drift

Displaces the named brand at turn two, every time, then progressively replaces brand engagement with generic category education — in extreme cases, how to make the product yourself. Rich, AI-legible brand content suppresses the drift.

Perplexity

Advocacy without a close

Builds the most thorough brand journeys, then stalls in a personalisation loop. Its live web retrieval can also inject newly-published competitors mid-conversation — a distinctive, real-time displacement risk.

Grok

Deterministic lock-in

Zero variance: the brand that owns a category’s primary criterion wins every run; the brand that concedes it loses every run. The highest purchase-stage steering of any model tested.

What it costs

The consideration gap is measurable

The consideration gap is the space between the brands an AI includes in its reasoning and the brands a consumer ultimately considers. It creates two compounding losses: consideration-set exclusion — a brand absent from the AI journey is absent from every channel downstream — and preference hardening, because a consumer who receives a detailed, multi-turn AI recommendation arrives at purchase far more committed than one who saw a single ad.

To size it, the research proposes a Revenue at Risk framework:

Revenue at Risk Annual sales × Discovery share (0.3–0.5) × Visibility gap × LLM share (0.1–0.25) × Conservatism factor (0.4)

Worked example: a brand with £100m revenue, a 0.4 discovery share, a 0.6 visibility gap and a 0.15 LLM share carries roughly £1.44m of AI-mediated revenue exposure a year — rising toward McKinsey’s 2028 projection as AI’s share grows. At £500m revenue and a 0.8 gap, the figure exceeds £9.6m.

The advertising trap

Why a sponsored placement can’t save you

Search advertising works by intercepting a consumer before their preference is formed. Inside an AI conversation, the preference is formed during the reasoning chain — so a sponsored placement appearing alongside the answer arrives after the decision has already happened. Paid placement without reasoning-chain presence amplifies absence, not presence.

The intervention with the strongest evidence base isn’t a traditional media buy — it’s citation architecture: systematically building your brand’s presence across the sources AI models actually draw on when they reason. It operates inside the reasoning chain rather than beside it, and it’s the prerequisite for every other AI investment, including the emerging commercial-handoff products like in-chat checkout.

The response

Three levers, one measurement stack

Persistent presence across the reasoning chain — arriving at the commercial handoff moment as the primary subject — is achievable, through three connected levers:

Lever 01

Citation architecture

Build generic consideration-set presence in the high-authority sources AI reasons from — so you’re on the shortlist before evaluation begins.

Lever 02

Displacement delay

Extend primary brand presence past the optimisation turn to the commercial handoff moment, using brand-specific, AI-legible content.

Lever 03

Platform-specific remediation

Diagnose and treat each platform’s distinct failure mode — because one fix does not transfer across four different mechanisms.

None of it is manageable without measurement built for the reasoning chain rather than the snapshot. That’s the foundation our companion work sets out in the PSOS™ methodology — an open, auditable KPI for AI visibility. You cannot manage what you cannot see; this research provides the evidence, and the measurement architecture required to see it.

The full research

Read the complete white paper

This article is an overview. The full paper details all three research phases, the four-instrument measurement stack, the fourteen-type displacement taxonomy, platform behaviour profiles and the complete methodology.

Citation: Sheals, P. (2026). The AI Consideration Gap: Structural Brand Displacement in AI Decision-Making. AIVO Standard. Zenodo. https://doi.org/10.5281/zenodo.19519860 · Licensed CC-BY-NC-4.0.