Three tools, the same prompts, three different answers
The AI-visibility category has a measurement credibility problem. It is not going to be solved by a fourth dashboard.
Run the same fifty prompts through three AI-visibility tools in the same week and you will get three materially different pictures of the same brand. This is now the most common complaint about the category, and the most common response to it — “our data is better” — is not an answer. It is the same claim the other two are making.
The disagreement is real, and most of it is explicable. Understanding where it comes from is the difference between a tool you can defend in a client meeting and one you cannot.
Four places the numbers diverge
The instrument. Some tools call a grounded model API. Some scrape the consumer interface. These are not the same measurement. Scraped consumer answers average around 743 words and 16 sources; API answers average around 406 words and 7. Roughly a quarter of API runs return no sources at all, and on identical prompts the API triggers a web search only about 42% of the time. A brand can be genuinely present in one instrument and genuinely absent in the other, with neither tool malfunctioning.
The sample size. A single run is one draw from a distribution, not a score. Day-to-day overlap in cited sources for identical prompts sits between 34% and 42%. If one tool runs each prompt once and another runs it three times, they will disagree, and the one running it once will disagree with *itself* next week.
The denominator. Share of voice is a share of a competitor set somebody chose. Two tools with different competitor lists are computing different quantities and calling them the same word.
The geography. Injecting “in South Africa” into a prompt is not the same as issuing the request from South Africa. One is a phrasing change that the model may or may not honour; the other is a parameter. Tools that do the first and describe it as localisation are producing a number about a different question.
Why “our data is better” cannot settle it
Because the claim is unfalsifiable from outside. Every tool in this market presents aggregates. Almost none let you open the observation behind one. Without that, a buyer comparing two tools is comparing two assertions, and will reasonably pick the cheaper assertion.
No metric is auditable without access to the observations behind it.
That line is becoming the standard test independent reviewers apply, and it is the right one. It is also uncomfortable, because it implies the thing most dashboards are selling — a clean number — is the part that cannot be checked.
What we do instead
- Declare the instrument. Every figure records how it was acquired, from where, in which language, at which model version, and whether a web search actually fired on that run.
- Store every answer. Verbatim, with its provenance, reachable from every metric it fed.
- Report ranges, not scores.
28–36% · n=150, never32. - Publish the prompt set. The denominator is printed with the report.
None of this makes our numbers agree with a competitor's. It makes the disagreement diagnosable — which is the only version of this argument an agency can win in front of a client.
The practical test to run on any vendor, including us: pick one figure in the interface and ask to see the answers behind it. Count how many clicks it takes.
See it against your own client list
A working demo runs your prompts, in your market, on live engines — not a sandbox with seeded data. Bring one client brand and three competitors.