Discovery is moving from the search box to the AI assistant. When a customer asks ChatGPT or Gemini for a recommendation, there is no page two, no list of ten blue links, no click-through rate to optimise. Your brand is either named in the answer — or it is invisible in the exact moment a decision is made.
That single shift breaks the tools we have relied on for twenty years. Impressions, rankings, backlinks and domain authority were all built for search engines. They say nothing about whether an AI model recommends you, omits you, or quietly hands the moment to a competitor.
PSOS™ — the Prompt-Space Occupancy Score — was created to close that gap: the first open, standardised way to measure how much space your brand occupies in the AI recommendation layer, with the audit trail and statistical rigour a boardroom can trust.
Three risks you can’t see without a measure
Vendor dashboards give partial snapshots, but they lack transparency, auditability and statistical confidence. That leaves leadership exposed on three fronts:
Governance blind spots. Boards cannot see or verify how their brand is being represented in AI recommendations.
Strategic misallocation. Teams keep over-investing in traditional SEO while under-investing in AI visibility.
Reputation risk. Misinformation, outdated data, or a competitor displacing you inside the answer goes undetected.
A single, board-grade number
PSOS distils the messy reality of AI recommendations into one composite score from 0 to 100. Just as PageRank gave the early web a shared measure of authority, and GAAP gave finance a shared standard for reporting, PSOS gives AI visibility a measure that is open, versioned and governed — designed to withstand scrutiny from executives, regulators and investors.
Crucially, it is not another black-box vendor metric. Every prompt set, weighting and formula is documented and versioned, so any score can be reproduced and defended.
Five dimensions of AI visibility
A brand’s presence in AI answers is more than a single mention. PSOS measures five distinct, auditable dimensions and combines them into the composite score.
Breadth
The share of relevant prompts where your brand appears — weighted by prominence, so being named first counts for more than being named last.
Depth
Whether that visibility persists over time. Measured across 30-, 60- and 90-day windows so a score reflects durable recall, not a campaign spike.
Resilience
Consistency across engines — ChatGPT, Gemini, Claude, Perplexity, Grok and more — weighted by market share, so you’re not reliant on a single model.
Sentiment
The tone of how you’re mentioned, from NLP polarity analysis, applied as a ±10% overlay. Presence matters; the quality of that presence matters too.
Decay
How fast visibility erodes without reinforcement — revealing which brands have naturally durable recall and how much investment is needed to hold position.
Two modes, one standard
PSOS is measured in two complementary ways, so the same standard fits both a global enterprise and a local business.
Test the models directly
- Curated clusters of 30–75 real discovery prompts per market, versioned and frozen for auditability.
- Run across the major engines, with each prompt repeated (3+ replicates) to cancel out AI randomness.
- Every response parsed for mentions, position, sentiment and refusals — logged with timestamps and provenance.
- Produces a composite with 95% confidence intervals and the five sub-scores reported separately.
Measure the signals models read
- For smaller organisations, uses proxy signals: citation diversity, review volume, structured-data freshness and offline credibility.
- A Confidence Index qualifies how reliable the result is when direct prompt-testing isn’t feasible.
- A cost-effective benchmark that still maps to the same 0–100 scale.
Local scores are read against three plain-English bands:
Built for governance, not just dashboards
Several tools monitor prompts or track brands. They fall short where it counts for a board: transparency, auditability and defensibility.
| Dimension | Typical vendor dashboard | PSOS™ |
|---|---|---|
| Methodology | Proprietary black box | Open, versioned, auditable |
| Audit trail | Limited or none | Full logs, frozen prompt sets, provenance |
| Statistical confidence | Not provided | Confidence intervals & error bands |
| Platform coverage | Single-engine focus | Aggregated across ChatGPT, Gemini, Claude and more |
| Governance readiness | Tactical reporting | Board-ready KPI, attributable to ROI |
Governance is designed in, not bolted on: prompt clusters are frozen once published, formulas are versioned (e.g. psos_v1.0.0), engine weightings are provenance-logged, and every affirmative mention is checked against a Citable Reference Unit so hallucinated citations are filtered out rather than counted.
The attribution layer
A score only matters if it moves the business. PSOS includes an attribution layer that links visibility changes to outcomes — leads, qualified opportunities and revenue — using difference-in-differences analysis around specific interventions, such as deploying structured data or launching a review campaign. That turns AI visibility from a marketing abstraction into a KPI a board can allocate budget against.
Validated at scale
PSOS was first applied in the AIVO 100™ Global Index — a study measuring AI visibility across more than 100,000 prompts, eight sectors and six leading AI platforms.
The score also tracks commercial reality. Early evidence shows PSOS correlating with aided brand recall (r = 0.67), organic referral traffic (r = 0.52) and marketing-qualified leads (r = 0.43) — positioning it as a leading indicator of performance, not just a visibility gauge. Sector patterns were revealing too: technology brands led on breadth but were volatile, while some healthcare brands carried misinformation risk, with outdated data surfacing in a meaningful share of prompts.
Known limitations
Part of being governance-grade is being honest about boundaries. Version 1.0 states its own:
Boundary conditions
- English-language dominance can skew global rankings toward US/UK brands; regional weighting is planned.
- Current scope excludes voice assistants (Siri, Alexa) and regional LLMs (Baidu, Yandex) due to access limits.
- The ±10% sentiment band may understate impact in reputation-sensitive sectors like healthcare and financial services.
- Standard 90-day windows can under-represent long-term equity in stable categories such as luxury and industrial B2B.
Read the complete PSOS™ Methodology
This article is an overview. The full v1.0 methodology — governance framework, both scoring modes, the six-tool architecture, validation evidence and certification pathways — is published open-access on Zenodo with a permanent DOI.
Citation: Sheals, P. (2025). PSOS™ Methodology v1.0 — Prompt-Space Occupancy Score (AI Visibility KPI). AIVO Standard. Zenodo. https://doi.org/10.5281/zenodo.17081529 · Licensed CC-BY-4.0.