The AI Visibility Benchmark: what we will publish, and what it cannot prove
A reproducible protocol for measuring how AI assistants describe brands, with pinned prompts, model snapshots, raw answers, and no invented causality.
An AI visibility number is useful only when someone else can reproduce it.
That sounds obvious. It is not. A dashboard can show a percentage while hiding the prompts, the model version, the number of runs, the search settings, and the rule used to count a mention. Without those details, the number is a story with arithmetic attached.
What The Model Says is preparing a benchmark built around the opposite standard: publish the inputs, the raw answers, the evaluation rule, and the limits alongside every result.
What the benchmark measures
The unit is an answer to a question a real buyer might ask an AI assistant. It is not a keyword ranking and not a count of pages indexed by a crawler. The first measurement will use a versioned prompt matrix covering four shapes:
- direct recommendations;
- comparisons with a named competitor;
- category questions;
- problem-first questions where the category is never named.
The planned first drop will use four pinned model snapshots: ChatGPT, Claude, Qwen, and DeepSeek. The exact provider identifiers, prompt count, repetition count, locale, and retrieval settings will be published in the run manifest before results are reported. This is a collection plan, not a benchmark result. The comparison slice will use a peer roster frozen before collection; it will not be chosen after looking at the answers.
The last shape matters. A brand can rank for its own category phrase and still be absent when a buyer describes the problem in ordinary language. The difference is one of the reasons a single visibility score is too blunt.
Our existing measurement guide explains the basic arithmetic and the ways a measurement can lie. The benchmark turns those rules into a repeatable release.
What every release will contain
Each drop will publish:
- the prompt-matrix version and the complete prompt list;
- the model provider and pinned snapshot identifier;
- the number of runs for every prompt;
- the raw answers in machine-readable formats;
- the evaluation criteria for a mention, citation, recommendation, and comparison;
- both the integer numerator and denominator behind every reported rate;
- a changelog explaining what changed since the previous drop.
Prompts will not be silently edited. If a prompt becomes misleading, it gets a new identifier and a new matrix version. Otherwise a time series can appear stable while its measuring instrument has quietly changed.
The public data page will be the index for these releases. The full limits and the no-causation rule live in the methodology, not in a footnote added after the numbers look interesting.
What the benchmark cannot prove
An increase in mentions does not prove that a particular article, technical change, or campaign caused the increase.
Answers are non-deterministic. Results can vary by account, locale, search setting, retrieval state, and model snapshot. A prompt matrix is also a choice made by the researcher; it can favour the positioning it was designed around.
So the benchmark will report movement, not invented attribution. A result can tell us what the chosen prompt set produced under the recorded conditions. It cannot tell us that one intervention caused the change unless a separate design supports that conclusion.
That distinction is the point of the project. A transparent limitation is more useful than a precise-looking number nobody can audit.
When the first drop is ready
This page is a protocol, not a result. The first dataset will be published only when the matrix, model snapshots, raw answers, and evaluation criteria are all recorded together.
Until then, the honest status is simple: the benchmark is in preparation. We would rather publish that sentence than fill the gap with synthetic numbers.