Primary Research · Retrieval Dynamics

Being found by AI isn’t the same as staying found.

100,000+ prompts reveal how AI assistants actually retrieve, rank and forget brands — why visibility built on the wrong sources swings 40–60% month to month, and what makes it durable.

Author: Paul Sheals · AIVO Standard™ 100,000+ prompts · all major LLMs Published Sept 2025 ≈ 8 min read

A brand can appear in an AI assistant’s answer this month and vanish the next — without doing anything wrong. Unlike traditional search, where rankings drift slowly, visibility inside large language models can swing dramatically from one month to the next.

Across more than 100,000 reverse-engineered prompts, spanning every major industry and model, one pattern is unmistakable: where your visibility is anchored decides whether it lasts. Brands leaning on the wrong sources aren’t just ranked lower — they’re structurally unstable.

“Companies relying primarily on Tier 2 and Tier 3 citations see their AI visibility fluctuate by 40–60% month on month.”
The framework

Not all citations are equal

AI visibility is tiered. The stability of your discoverability depends on which layer of sources you’re anchored in — and each tier plays a different role.

Tier 1 · the anchor

Foundational citations

Tie-breaker & durability.
  • Wikidata, Wikipedia, GitHub, Hugging Face, Zenodo, schema.org
  • Canonical, evidence-based — ingested into training data & knowledge graphs
  • When competition is high, Tier 1 presence wins the shortlist
  • Slow to decay, hard to displace
Tier 2 · the validator

Industry validation

Multiplies odds of inclusion.
  • Crunchbase, G2, analyst reports, trusted press, trade databases
  • Contextual credibility — “formally recognised in its sector”
  • Strengthens presence, but not durable on its own
Tier 3 · the entry ticket

Topical & recency

Gets you in the door.
  • Websites, blogs, reviews, news, social chatter
  • Fast-moving and easy to access — ensures initial inclusion
  • Least stable: mentions rise and fall, driving the volatility

In practice: if two brands are mentioned equally across Tier 3 and Tier 2, the one with Tier 1 citations wins the shortlist — consistently.

The finding

The cost of a fragile foundation

The headline result of the research programme is the sheer instability of visibility that isn’t anchored:

40–60%
month-on-month visibility swing for brands relying primarily on Tier 2/3 sources.
100,000+
reverse-engineered prompts across industries and every major LLM, over 12+ months of R&D.
3–5
entities in a typical AI shortlist — a narrow cut where durability decides who survives.

Brands with strong Tier 1 anchoring remain stable across model refreshes and months of querying. Tier 1 doesn’t grant immunity — models still rebalance their training data — but it delivers orders of magnitude more durability than Tier 2/3 reliance.

Under the hood

How an AI builds its shortlist

LLMs don’t “search” the web like Google. They blend static training knowledge, live retrieval and ranking heuristics — and the research maps five consistent stages.

1

Candidate retrieval

A broad, Tier 3-heavy longlist of 10–20 entities is pulled from websites, news and recent content.

2

Deduplication & clustering

Variants of the same brand are collapsed into one entity; frequency across sources adds early weight.

3

Authority weighting

Candidates are cross-checked against Tier 2 and Tier 1. Entities validated across multiple tiers gain authority.

4

Relevance & semantic alignment

The model checks which candidates best match the user’s actual intent — authority matters less if relevance is missing.

5

Shortlist curation the cut

The final 3–5, scored on a blend of coverage, authority and relevance/recency. This is where durability wins or loses.

Platform intelligence

Every model weights the tiers differently

The same brand can be stable on one platform and volatile on another, because each model biases the tiers its own way.

ChatGPT

Tier 1-anchored

Leans on canonical sources — Wikipedia, Wikidata, schema.org, DOIs. More stable over time, but can lag on recency.

Gemini

The hybrid

Balances Tier 1 canonical authority with Tier 3 recency, using Google’s indexing to blend durability and freshness.

Perplexity

Recency-heavy

Prefers Tier 3 freshness validated by Tier 2; confidence rises when an entity appears across multiple tiers.

Grok

Chatter-biased

Social-relevance weighting — rewards entities “in the conversation” on X and topical news, even with weaker Tier 1.

Retrieval inverts for time-sensitive queries. Canonical questions (“best CRM platforms?”) start at Tier 1 and work down; time-sensitive ones (“cheap holiday next week”) start at Tier 3 and work up. That’s why small local players surface briefly for hot queries — but rarely with any durability.

The governance gap

Most tools measure the wrong thing

Current AI-visibility tools track where a brand appears right now — mostly Tier 3 mentions. That’s useful for tactical monitoring, but it says nothing about durability, decay, risk or causality. The danger is a false sense of success: a temporary Tier 3 spike gets reported as an “improvement,” and budget follows a number that was never stable.

Closing that gap needs a governance-grade KPI that weights presence across all three tiers and accounts for decay — which is exactly what the Prompt-Space Occupancy Score (PSOS™) was built to be: a stable, comparable, auditable measure boards, investors and regulators can trust.

The full paper

Read the complete white paper

This article is an overview. The full paper details the tiered citation framework, the five-stage retrieval process, model-by-model behaviour, the volatility & decay analysis, worked case examples and the PSOS methodology — published open-access on Zenodo with a permanent DOI.

Citation: Sheals, P. (2025). AI Visibility Retrieval Dynamics: Tiered Citation Platforms, Decay, and the Governance Gap in LLM Discoverability. AIVO Standard. Zenodo. https://doi.org/10.5281/zenodo.17117353 · Licensed CC-BY-4.0.