This page exists so you can check our work. It states what we actually ask, how many times we ask it, how wide the uncertainty is, and — the part most tools leave out — what we cannot see at all.
If a figure isn't on this page, we don't publish it.
Every visibility figure we show comes from an answer a model actually gave to a question we actually asked, stored with its full text so it can be re-read later. There is no extrapolation from clickstream data, no modelled search volume, no inferred "AI rank".
The cost of that rule is that our numbers are sometimes smaller, and sometimes absent. A brand with too few answers to be sure about shows not enough data rather than a confident-looking percentage. We think a visible gap is more useful than a number you can't lean on.
A probe is one prompt, sent to one model, once. The answer is scored for three things: whether the brand was named, how favourably it was described, and which sources the model cited. The full response text is stored alongside the score.
Nothing is aggregated across models. Because two engines agree on which brands to name far less often than people assume, a single blended "visibility score" hides the thing you need to act on. Every figure breaks out per engine, always.
| Provider | Model probed | We call it | Answers collected | Most recent |
|---|---|---|---|---|
| OpenAI | gpt-4o-mini | ChatGPT | 4,809 | 20 Aug 2026 |
| Anthropic | claude-haiku-4-5-20251001 | Claude | 4,492 | 19 Aug 2026 |
| gemini-2.5-flash-lite | Gemini | 2,873 | 20 Aug 2026 |
These are the exact model identifiers, not marketing names. "Claude" here means Claude Haiku 4.5 through the API — not the Claude consumer app, which has its own system prompt and retrieval behaviour and can answer differently. Web search is off by default (it is opt-in per run) and is recorded on every answer, because it materially changes what a model cites.
| What | Count | Read it as |
|---|---|---|
| Answers collected | 12,375 | Every stored answer, all engines, since we started. Nothing is estimated into this number. |
| Period | 27 May – 20 Aug 2026 | An 86-day span, but answers were collected on 33 separate days inside it. We probe in rounds, not continuously, and we do not call that "daily". |
| Engines | 3 | ChatGPT, Claude and Gemini. What we do not read is section 06. |
| Brands | 23 | Each with its own prompt set and competitor list. Three of them are industry-study rosters rather than client brands. |
| Prompts probed | 1,418 | Distinct questions asked at least once. 2,771 are configured; we count the ones that have actually produced an answer. |
| Markets | 9 | Country-tagged prompt sets. What "measured in a market" is worth depends on the language the answers came back in — section 05. |
| Citations recorded | 31,667 | Source URLs the engines themselves returned. |
| Answers with web search on | 3,755 | 30% of the corpus. Search is opt-in per run and costs roughly 17× an ungrounded answer, so most rounds are ungrounded. Grounded and ungrounded answers are never averaged together. |
This is a small corpus next to the industry studies we cite elsewhere, and we would rather say so than pad it. Every one of these answers exists as stored text.
Claude answers stopped on 19 August 2026. Our own Anthropic account reached the spend limit set on it. That limit has since lifted — access was re-tested and confirmed working on 1 September 2026 — but no Claude round has run yet, so any Claude figure you see still comes from answers collected on or before 19 August, and rounds run since then are two-engine rounds. This was our billing rather than a fault in the product, and the remaining gap is the data, not the access — but it changes what a number means, which is why it is on this page.
Language models are not deterministic. Ask the same question twice and you can get two different brand lists. That means a mention rate is a sample statistic, and a sample statistic without an interval is a guess wearing a suit.
We use the Wilson score interval at 95% (z = 1.96). It's chosen over the textbook normal approximation because it stays sensible at the edges — at 0 hits out of 12, or at small n, where the naive formula produces intervals that run below zero.
| Rule | Value | Why |
|---|---|---|
| Minimum sample | n ≥ 7 | Below 7 answers we publish no rate at all — the cell reads "not enough data". A small n is not a zero; it means we haven't looked hard enough yet. |
| Interval | Wilson, 95% | Shown next to every rate. Two rates whose intervals overlap are not a ranking, and we don't present them as one. |
| Precision | ±10pp ≈ 100 runs | Roughly what it takes to pin a mid-range rate to ten points. Tighter answers need more runs, and more runs cost money — we'd rather tell you the width than hide it. |
Measured counted directly from stored answers · Derived arithmetic on measured values, no new assumptions · Not offered we don't publish it
| Metric | Basis | How it's computed |
|---|---|---|
| Mention rate | Measured | Answers naming the brand ÷ answers collected, per engine, with a Wilson interval. Branded prompts (where the brand is named in the question) are reported separately — averaging them in is how a tool tells you 62% when the honest earned figure is 0%. |
| Sentiment | Measured | Scored from the answer text where the brand appears, and only on answers that actually mention the brand. The scorer reads English; on an answer it cannot read it abstains rather than returning neutral, because a 0.0 nobody wrote is a fabricated data point. That currently silences 111 of 4,836 brand-mentioning answers (2.3%) — 0 of 3,757 English ones, but 27 of 32 Polish and 22 of 39 Portuguese. Sentiment is therefore thinnest exactly where a market was answered in its own language. |
| Citations | Measured | URLs the model itself returned. We never infer a citation from the fact that a page ranks in Google. |
| Share of voice | Derived | Your mentions ÷ all tracked-brand mentions on the same prompt set. Only comparable within one prompt set and one engine — it is not a market share. |
| Website readiness | Derived | Rule checks against pages we could actually fetch. A site we could not read returns no score at all — the field is left empty and labelled unread, never given a low grade. A bot wall tells you nothing about the content behind it, and grading it anyway would invent a finding. |
| Prompt volume | Not offered | Nobody can currently observe how often real people ask an assistant a given question. Tools that publish this figure derive it from clickstream panels and say so in their own documentation. We'd rather leave the field blank. |
| "AI rank" | Not offered | There is no ranked results page to hold a position in. We report how often you're named, not what position you hold, because the second thing doesn't exist. |
A market is a country. A country is not a language, and an assistant asked a localized question will often answer in English anyway. If nobody checks, an English answer counted into a Polish market's number turns into "your Polish visibility is 38%" — a sentence about Poland built from answers that were never Polish.
So every stored answer now carries the language it was actually written in, detected from the answer text rather than assumed from the question. Where a market's answers came back in another language, the number on screen says so, next to itself, with the share. It is a statement of what was measured, not a warning: these are real answers that really mention the brand — the only thing that was ever wrong is what we called them.
Once the same question has been re-asked on the same engine and answered in the market's own language, the earlier off-language answer is excluded from that market's figure and stays in the database. Same question and same engine, deliberately: dropping every off-language answer the moment any local one exists would shrink some markets to one or two answers, and a "market figure" computed from one answer hides that we measured at all.
| Market | Answers | In a local language | What that means |
|---|---|---|---|
| United States | 121 | 100% | English is the right answer here, not a fallback. These markets are measured as intended. |
| United Kingdom | 56 | 100% | |
| Spain | 174 | 100% | Fully measured in Spanish. |
| Italy | 174 | 100% | Fully measured in Italian. |
| France | 1,881 | 88% | The residual is older answers, before the fix. Every French round since 2 July 2026 has come back in French. |
| Germany | 54 | 41% | Majority English. The German figure is labelled accordingly until it is re-probed. |
| Poland | 232 | 37% | Re-probing has started: 85 Polish answers now exist where there were almost none. |
| Portugal | 303 | 32% | Same — partly re-probed, and the label moves on its own as it is. |
| Switzerland | 158 | 20% | The weakest market we have. Switzerland has three national languages, so we refuse to guess one: a Swiss prompt set has to name its language or it does not run. |
Across all markets: 2,418 of 3,153 market-tagged answers (77%) came back in a language of their market. Language was detected on 12,297 of our 12,375 answers; the remaining 78 were too short or too mixed to call, and are recorded as unknown rather than assigned a language.
No AI-visibility tool covers everything, ours included. These are the gaps we know about. We'd rather you learn them here than discover them after signing.
| Surface | Status | Why |
|---|---|---|
| Meta AI | Not measured | Very large consumer reach, no API that permits measurement. This is a genuine blind spot for the whole category, not just for us. |
| Google AI Overviews / AI Mode | Not measured | Distinct surfaces from the Gemini API, with different retrieval. We don't treat a Gemini API answer as a proxy for either. Both were included in the 18 August 2026 pilot at 20 answers each; that is one day, not a measurement, and it feeds nothing. |
| Perplexity, Grok, Mistral, DeepSeek, Llama | Built, piloted once, not tracked | The connectors work and we ran them once, on 18 August 2026: 81 answers from Perplexity and 20 each from the others. That is a single day and it has not been repeated, so nothing here is a trend and none of it reaches a client's numbers. When one starts running on a cadence it will be added to the engine table in section 02, with its model identifier and its count, before it is claimed anywhere else. |
| Microsoft Copilot | Never probed | No connector, no pilot, no stored answers. We have nothing to say about it. |
| Consumer apps vs APIs | Known divergence | We measure APIs. The consumer apps add system prompts, memory and live retrieval, so their answers can differ from ours. We don't claim otherwise. |
| Model drift | Uncontrolled | Providers update models without notice. A change in your numbers can be a change in the model rather than a change in your visibility, which is why we record the exact model identifier on every stored answer. |
Run the same brand through two AI-visibility tools and you will get two different answers. That isn't one of them lying. It's four things stacking up:
If our number is lower than another tool's, the useful question is which prompt set each of us asked, and how many times. We'll show you ours.
The rules behind our recommendations are versioned and graded by evidence strength, from peer-reviewed research down to unverified industry folklore. When a rule is corrected or retired, the change is recorded with its date and its source rather than quietly edited — including the cases where we were previously wrong.
That ledger also runs in reverse: we keep an explicit list of widely-sold tactics the evidence does not support, so a recommendation can be refused as well as made.