A brand can appear in an AI assistant’s answer this month and vanish the next — without doing anything wrong. Unlike traditional search, where rankings drift slowly, visibility inside large language models can swing dramatically from one month to the next.
Across more than 100,000 reverse-engineered prompts, spanning every major industry and model, one pattern is unmistakable: where your visibility is anchored decides whether it lasts. Brands leaning on the wrong sources aren’t just ranked lower — they’re structurally unstable.
Not all citations are equal
AI visibility is tiered. The stability of your discoverability depends on which layer of sources you’re anchored in — and each tier plays a different role.
Foundational citations
- Wikidata, Wikipedia, GitHub, Hugging Face, Zenodo, schema.org
- Canonical, evidence-based — ingested into training data & knowledge graphs
- When competition is high, Tier 1 presence wins the shortlist
- Slow to decay, hard to displace
Industry validation
- Crunchbase, G2, analyst reports, trusted press, trade databases
- Contextual credibility — “formally recognised in its sector”
- Strengthens presence, but not durable on its own
Topical & recency
- Websites, blogs, reviews, news, social chatter
- Fast-moving and easy to access — ensures initial inclusion
- Least stable: mentions rise and fall, driving the volatility
In practice: if two brands are mentioned equally across Tier 3 and Tier 2, the one with Tier 1 citations wins the shortlist — consistently.
The cost of a fragile foundation
The headline result of the research programme is the sheer instability of visibility that isn’t anchored:
Brands with strong Tier 1 anchoring remain stable across model refreshes and months of querying. Tier 1 doesn’t grant immunity — models still rebalance their training data — but it delivers orders of magnitude more durability than Tier 2/3 reliance.
How an AI builds its shortlist
LLMs don’t “search” the web like Google. They blend static training knowledge, live retrieval and ranking heuristics — and the research maps five consistent stages.
Candidate retrieval
A broad, Tier 3-heavy longlist of 10–20 entities is pulled from websites, news and recent content.
Deduplication & clustering
Variants of the same brand are collapsed into one entity; frequency across sources adds early weight.
Authority weighting
Candidates are cross-checked against Tier 2 and Tier 1. Entities validated across multiple tiers gain authority.
Relevance & semantic alignment
The model checks which candidates best match the user’s actual intent — authority matters less if relevance is missing.
Shortlist curation the cut
The final 3–5, scored on a blend of coverage, authority and relevance/recency. This is where durability wins or loses.
Every model weights the tiers differently
The same brand can be stable on one platform and volatile on another, because each model biases the tiers its own way.
Tier 1-anchored
Leans on canonical sources — Wikipedia, Wikidata, schema.org, DOIs. More stable over time, but can lag on recency.
The hybrid
Balances Tier 1 canonical authority with Tier 3 recency, using Google’s indexing to blend durability and freshness.
Recency-heavy
Prefers Tier 3 freshness validated by Tier 2; confidence rises when an entity appears across multiple tiers.
Chatter-biased
Social-relevance weighting — rewards entities “in the conversation” on X and topical news, even with weaker Tier 1.
Retrieval inverts for time-sensitive queries. Canonical questions (“best CRM platforms?”) start at Tier 1 and work down; time-sensitive ones (“cheap holiday next week”) start at Tier 3 and work up. That’s why small local players surface briefly for hot queries — but rarely with any durability.
Most tools measure the wrong thing
Current AI-visibility tools track where a brand appears right now — mostly Tier 3 mentions. That’s useful for tactical monitoring, but it says nothing about durability, decay, risk or causality. The danger is a false sense of success: a temporary Tier 3 spike gets reported as an “improvement,” and budget follows a number that was never stable.
Closing that gap needs a governance-grade KPI that weights presence across all three tiers and accounts for decay — which is exactly what the Prompt-Space Occupancy Score (PSOS™) was built to be: a stable, comparable, auditable measure boards, investors and regulators can trust.
Where this sits
Trust Resilience
The tiered-trust architecture this evidence underpins — how to engineer durable, layered trust that survives decay.
PSOS™ — Prompt-Space Occupancy Score
The governance-grade KPI that weights visibility across all three citation tiers and accounts for decay.
Read the complete white paper
This article is an overview. The full paper details the tiered citation framework, the five-stage retrieval process, model-by-model behaviour, the volatility & decay analysis, worked case examples and the PSOS methodology — published open-access on Zenodo with a permanent DOI.
Citation: Sheals, P. (2025). AI Visibility Retrieval Dynamics: Tiered Citation Platforms, Decay, and the Governance Gap in LLM Discoverability. AIVO Standard. Zenodo. https://doi.org/10.5281/zenodo.17117353 · Licensed CC-BY-4.0.