Primary Research · Product Data

Not invisible. Illegible.

Why AI can’t recommend what it can’t read — the Product Data Legibility Gap.

AI doesn’t read your master product file. It reads a fragmented, contradictory version of it scattered across every retailer — and when your data can’t be read cleanly, the recommendation goes to whoever’s can.

By Tim de Rosen & Paul Sheals AIVO Meridian · WP-2026-15 Published May 2026 ≈ 7 min read

When a shopper asks an AI which foundation suits combination skin, the model doesn’t open a retailer’s website and read your product page. It draws on what it ingested during training and retrieval: Amazon listings, Sephora and Ulta pages, your own DTC site, beauty editorial, reviews. Every one of them describes the same product with different terminology, different attribute structures, different shade conventions.

When those sources conflict, the model doesn’t average them. It preferences the source it has the highest confidence in — typically the biggest retailer with the most densely populated record — and treats that version as the truth, whether or not it matches your canonical positioning. The result is that the AI’s version of your product is not your version. It’s a probabilistic reconstruction assembled from whatever its sources happened to contain.

“The brand was not invisible. It was illegible.”
The distinction

Being seen isn’t being read

This shift has been treated as a visibility problem — measure how often you’re cited, produce content to get cited more. But a brand can be highly visible in AI outputs and still lose the final purchase recommendation to a competitor, or even to its own inconsistently-indexed SKU. Citation is necessary but not sufficient.

What decides whether a brand wins the recommendation isn’t how often the model has seen it. It’s how clearly and consistently the model can read it. That delta — between what a brand intends its product to be understood as, and what the AI can actually read and reason from at the decision moment — is the Product Data Legibility Gap.

The finding

Twelve attributes decide it

Across a composite dataset of ~800–1,000 cosmetics SKUs, probed across ChatGPT, Gemini, Perplexity and Grok, a Pareto pattern emerged:

12
PIM attributes explain roughly 80% of cross-retailer LLM citation variance — out of 300+ in a typical schema.
The cruel twist
the attributes AI cites most are the ones brands manage least consistently across retailers.
At the master
fragmentation starts in the brand’s own product file — then multiplies at every retailer.

The pattern is stark when you line up how often AI cites an attribute against how consistent it is across a brand’s retailers:

AttributeAI cites itConsistent across retailersOutcome impact
Product name format92%24%High
Skin type / use case88%31%High
Coverage / finish85%38%High
Key ingredients81%19%High
Shade name / code78%42%High
SPF / sun protection28%84%Low
Country of origin14%91%Low

The high-impact attributes are cited constantly and managed inconsistently. The consistently-managed ones barely matter. Brands are getting the priority exactly backwards.

The mechanism

It starts in your own file

Fragmentation doesn’t begin at the retailer — it begins in the brand’s own master product file and propagates downstream. In a representative 900-SKU portfolio, naming inconsistencies appear 15–30 times in the master itself before any retailer even touches the feed. Each retailer then applies its own ontology on top, so three already-inconsistent strings become four, five or six distinct product representations in the AI’s source diet.

And the AI doesn’t reconcile them. It adopts the dominant retailer’s version — including that retailer’s name string and copy — so the parent brand quietly loses ground to a retailer-specific SKU framing. The highest-revenue categories, the research found, consistently showed the lowest data consistency: the stakes are highest exactly where the data is most fragile.

Why it’s about to get worse

In agentic commerce, it’s a hard filter

Today an AI weighs conflicting data and still forms a recommendation. An agentic model — one executing a purchase on a consumer’s behalf — doesn’t weigh probabilities. It reads the attribute and acts.

“In an agentic commerce environment, PIM alignment is not a measurement problem. It is a distribution problem.”

A product that declares “Cruelty-Free: Yes” on one retailer but leaves the field blank on another will be hard-filtered out by an agent told to buy a cruelty-free product. It won’t infer; it finds the attribute absent and moves to the next result. At that point, data synchronisation across your retailer footprint stops being hygiene and becomes a gate on whether you can transact at all.

The remediation

Fixable — and measurable

Closing the gap needs neither more content nor better prompts. It needs alignment at the attribute level, anchored to the attributes the AI actually cites — done in three moves.

1

A canonical master record

One authoritative, de-duplicated definition of each product’s intended positioning — the version your master file should already have contained.

2

Per-retailer correction files

Field-level fixes for each retailer (typically 10–14 per product), pushed straight into the PIM platforms you already run — most operationalisable without any retailer negotiation.

3

Measure the delta

Baseline the recommendation rate before, re-probe after — the change is the proof. Any vendor can generate corrections; the measurement envelope is what proves they moved AI outcomes.

The full paper

Read the complete working paper

This article is an overview. The full paper details the diagnostic methodology, the full twelve-attribute hierarchy, the SKU-displacing-parent-brand pattern, the remediation architecture and the before/after measurement envelope — published open-access on Zenodo with a permanent DOI.

Citation: de Rosen, T. & Sheals, P. (2026). The Product Data Legibility Gap: Why LLMs Cannot Recommend What They Cannot Read. AIVO Meridian, Working Paper WP-2026-15. Zenodo. https://doi.org/10.5281/zenodo.20322459 · Licensed CC-BY-4.0.