Methods and evidence
How Caeliai measures AI shopping visibility.
Every number we publish comes from a defined test. We record shopping conversations with ChatGPT and Gemini, apply the same classification rules, and state what the result cannot tell you.
The measurement, step by step.
Real conversations, not APIs alone.
We ask ChatGPT and Gemini defined shopping questions in the consumer apps being studied. Where a study also reads catalog data, such as Shopify’s endpoints, the paper says so explicitly.
Repeated prompts.
AI answers change between runs. When the question is about a pattern rather than a single product, we run the same prompt multiple times and report rates across runs, not a single answer.
Record what appears.
For each answer, we record which brands are named, which products appear, whether a visual product card renders, and where each returned link leads.
Follow the returned link.
The core question: where does the returned link lead? Each observation is classified as owns the path (the brand’s own product page), leaks the path (a retailer, marketplace, or reseller), or no path (named with no usable link).
Check the details.
Is it the right product? Is the price current? Is the first seller the official one? We record wrong sellers and stale listings separately from the brand mention.
Capture the evidence.
We save observations as screenshots, response records, and structured data with dates. Published figures state their sample size and measurement date. We do not publish an undated chart.
Record the test context.
The record includes the prompt, whether the brand was named, the interface, session conditions, date, repeat policy, and any known personalization. One account is not treated as a universal view.
Results vary between runs.
No single AI answer is definitive. The same prompt can return different brands, products, and links at different times. That variation is a property of the systems being measured, so one screenshot cannot describe the full result.
Repetition separates a pattern from normal variation. A brand that appears in one run out of ten is a single observation. A brand that returns the same reseller across repeated runs has a pattern worth reviewing. Rates across repeated observations, with the N stated, are the unit of evidence. Single observations are labeled as single observations.
Limitations we state up front.
These systems change without notice, so every finding is timestamped and can go stale. Most findings are correlational. A catalog match alongside a product card is a useful observation, not proof of cause. Sample sizes are stated on every figure, and small ones are called early signals. Most evidence on this site comes from our measurements; the methods are documented so others can inspect or rerun them.
Disclosure.
Caeliai research is independent. No brand, platform, or vendor pays for placement in the research, and advisory clients do not buy better results in published studies. Where a study involves a client’s store, that is disclosed. Public research stays open and free to read.
Questions about the method, or a suspected error? Email contact@caeliai.com and Caeliai will respond directly. See also about Caeliai and a sample diagnosis built with this method.