This page exists so you can check our work. It states what we actually ask, how many times we ask it, how wide the uncertainty is, and — the part most tools leave out — what we cannot see at all.
If a figure isn't on this page, we don't publish it.
Every visibility figure we show comes from an answer a model actually gave to a question we actually asked, stored with its full text so it can be re-read later. There is no extrapolation from clickstream data, no modelled search volume, no inferred "AI rank".
The cost of that rule is that our numbers are sometimes smaller, and sometimes absent. A brand with too few answers to be sure about shows not enough data rather than a confident-looking percentage. We think a visible gap is more useful than a number you can't lean on.
A probe is one prompt, sent to one model, once. The answer is scored for three things: whether the brand was named, how favourably it was described, and which sources the model cited. The full response text is stored alongside the score.
Nothing is aggregated across models. Because two engines agree on which brands to name far less often than people assume, a single blended "visibility score" hides the thing you need to act on. Every figure breaks out per engine, always.
| Provider | Model probed | What we call it |
|---|---|---|
| Anthropic | claude-haiku-4-5 | Claude |
| OpenAI | gpt-4o-mini | GPT |
| gemini-2.5-flash-lite | Gemini |
These are the exact model identifiers, not marketing names. "Claude" here means Claude Haiku 4.5 through the API — not the Claude consumer app, which has its own system prompt and retrieval behaviour and can answer differently. Web search is off by default and is labelled on any run where it was on, because it materially changes what a model cites.
Language models are not deterministic. Ask the same question twice and you can get two different brand lists. That means a mention rate is a sample statistic, and a sample statistic without an interval is a guess wearing a suit.
We use the Wilson score interval at 95% (z = 1.96). It's chosen over the textbook normal approximation because it stays sensible at the edges — at 0 hits out of 12, or at small n, where the naive formula produces intervals that run below zero.
| Rule | Value | Why |
|---|---|---|
| Minimum sample | n ≥ 7 | Below 7 answers we publish no rate at all — the cell reads "not enough data". A small n is not a zero; it means we haven't looked hard enough yet. |
| Interval | Wilson, 95% | Shown next to every rate. Two rates whose intervals overlap are not a ranking, and we don't present them as one. |
| Precision | ±10pp ≈ 100 runs | Roughly what it takes to pin a mid-range rate to ten points. Tighter answers need more runs, and more runs cost money — we'd rather tell you the width than hide it. |
Measured counted directly from stored answers · Derived arithmetic on measured values, no new assumptions · Not offered we don't publish it
| Metric | Basis | How it's computed |
|---|---|---|
| Mention rate | Measured | Answers naming the brand ÷ answers collected, per engine, with a Wilson interval. Branded prompts (where the brand is named in the question) are reported separately — averaging them in is how a tool tells you 62% when the honest earned figure is 0%. |
| Sentiment | Measured | Scored from the answer text where the brand appears. Reported only on answers that actually mention the brand. |
| Citations | Measured | URLs the model itself returned. We never infer a citation from the fact that a page ranks in Google. |
| Share of voice | Derived | Your mentions ÷ all tracked-brand mentions on the same prompt set. Only comparable within one prompt set and one engine — it is not a market share. |
| Website readiness | Derived | Rule checks against pages we could actually fetch. Pages behind a bot wall are reported as unread, never scored as bad. A site we can't read is an unknown, not a failure. |
| Prompt volume | Not offered | Nobody can currently observe how often real people ask an assistant a given question. Tools that publish this figure derive it from clickstream panels and say so in their own documentation. We'd rather leave the field blank. |
| "AI rank" | Not offered | There is no ranked results page to hold a position in. We report how often you're named, not what position you hold, because the second thing doesn't exist. |
No AI-visibility tool covers everything, ours included. These are the gaps we know about. We'd rather you learn them here than discover them after signing.
| Surface | Status | Why |
|---|---|---|
| Meta AI | Not measured | Very large consumer reach, no API that permits measurement. This is a genuine blind spot for the whole category, not just for us. |
| Google AI Overviews / AI Mode | Not measured | Distinct surfaces from the Gemini API, with different retrieval. We don't treat a Gemini API answer as a proxy for either. |
| Perplexity, Copilot, Grok | Not measured | Not currently probed. When we add one, it will appear on this page with its own model identifier before it appears in the product. |
| Consumer apps vs APIs | Known divergence | We measure APIs. The consumer apps add system prompts, memory and live retrieval, so their answers can differ from ours. We don't claim otherwise. |
| Model drift | Uncontrolled | Providers update models without notice. A change in your numbers can be a change in the model rather than a change in your visibility, which is why we record the exact model identifier on every stored answer. |
Run the same brand through two AI-visibility tools and you will get two different answers. That isn't one of them lying. It's four things stacking up:
If our number is lower than another tool's, the useful question is which prompt set each of us asked, and how many times. We'll show you ours.
The rules behind our recommendations are versioned and graded by evidence strength, from peer-reviewed research down to unverified industry folklore. When a rule is corrected or retired, the change is recorded with its date and its source rather than quietly edited — including the cases where we were previously wrong.
That ledger also runs in reverse: we keep an explicit list of widely-sold tactics the evidence does not support, so a recommendation can be refused as well as made.