The declared instrument.
Every AI-visibility vendor produces numbers. Almost none will tell you exactly how. This page is our full disclosure: what we ask, how often, from where, with what, and what each figure can and cannot support. If a number on a client report is ever questioned, the answer is on this page.
A run is the only thing we actually observe.
One prompt, one engine, one language, one clean session, one moment in time. Everything else in the product — every rate, every share, every trend point — is an aggregate over runs, and every aggregate states how many runs are behind it.
A run produces an answer. That answer is written to storage, verbatim, with its provenance, before anything is derived from it. This ordering is deliberate: a figure computed from an answer nobody kept is a figure nobody can check.
Recorded on every capture
| Field | Pins down |
|---|---|
| method | Grounded API, or SERP vendor, and which one |
| engine + version | So a step-change can be attributed to a model update |
| locale + geo params | Where the question was asked from |
| language | Which phrasing, in which language |
| captured_at | When, to the second |
| search_triggered | Whether the engine actually performed a web search |
Three samples a week. Here is exactly what that buys.
This is the number every serious buyer should interrogate, so we will do it for you — including the finding that argues against us.
The finding against us
The strongest measurement study of 2026 found that stabilising a single prompt’s brand-detection estimate to a standard error under 0.10 takes roughly seven runs a day. Day-to-day overlap in cited sources for identical prompts sits at 34–42%.
Why it does not sink the design
Because nobody buys a report about one prompt. Fifty prompts × three samples is roughly 150 observations per engine per cycle, which supports an honest share-of-voice band of about ±7–8 points — provided it is reported as a band. Reported as a single figure, the same data is a lie with a decimal point.
And what we changed instead
Past around five repeats, an additional paraphrase or language reduces error 4–15× faster than another repeat of the same phrasing. So the answer to “three is not enough” is breadth, not a bigger N — and prompt groups and languages became schema rather than settings.
The honest summary: our per-prompt claims are directional and we label them as such with a stability class rather than a percentage. Our aggregate claims are properly powered and are reported with their intervals. A competitor reporting a single number off a comparable sample is making a stronger claim on weaker grounds.
Bands, run counts, stability classes.
What a band is
A binomial confidence interval over the period’s runs. Not a smoothing, not a min-and-max of daily values, and not a presentational flourish — the honest width of what this many observations can say about an underlying rate.
How to read two of them
- Bands that do not overlap: the movement is larger than sampling explains. Something happened.
- Bands that overlap: no detectable movement. The product says so rather than drawing an arrow, and the alerting system stays quiet.
- A band that widened: coverage dropped. An engine failed, or fewer runs completed. The figure did not get worse; our confidence in it did.
The third period is not a decline. It is a fortnight where two engine legs did not complete, so the band is wider and overlaps the one before it. Reporting a point estimate here would have invented a two-point drop out of a vendor outage.
What we will not claim: that this is what your customer sees.
This is the category’s biggest fault line and most vendors walk straight over it. We are going to be explicit about which side we are on and what it costs.
| Grounded API (what we use) | Consumer-interface scraping | |
|---|---|---|
| Fidelity to what a person sees | Lower — no router, no memory, no consumer system prompt | Higher |
| Typical answer length | ~406 words | ~743 words |
| Typical sources cited | ~7 (and ~25% of runs cite none) | ~16 |
| Web search actually fires | ~42% of runs on identical prompts | Always, by construction |
| Terms of service | Sanctioned, contractual | Grey — and your client inherits the exposure |
| Cost and scale | Viable at agency pricing | Expensive and brittle |
The claim we actually make
Not “our numbers are what consumers see” — nobody can defend that, including the scrapers, because consumer answers are personalised per user. What we claim is narrower and checkable:
Every number declares exactly how it was captured, and you can open the raw answers behind it.
What hardens it
search_triggeredon every grounded capture. Without it, “the engine looked and you were not there” and “the engine never looked” are the same data point. Nearly 58% of API runs in one study never searched.- Raw-answer click-through from every metric. The feature reviewers use to separate credible tools from dashboards.
- This page, versioned. When the instrument changes, it appears in the changelog with the date, so a moved number can always be attributed to the world or to us.
Location is a parameter, not a phrase.
Adding “in South Africa” to a prompt is a phrasing change the model may or may not honour, and the practice has been the subject of the sharpest published criticism of a market leader. We pass location as a request parameter wherever an engine supports one, and we record what was passed on every capture — including when the answer is “nothing, this engine has no location control”.
| Engine | Location control available | What we do |
|---|---|---|
| ChatGPT | Request-level user location | Pass ZA location parameters |
| Claude | Country, city, timezone on the search tool | Pass all three; validated against unlocated control runs |
| Gemini | None on the consumer developer API; coordinates via Vertex | Route through Vertex for geo-grounding, and label it |
| Perplexity | Full location filters | Pass ZA filters |
| Google AI Overviews | Vendor geo + device | ZA geo, desktop and mobile treated separately |
| Google AI Mode | Vendor geo + device | As above |
Language is treated as a first-class dimension rather than a filter, because it is the single largest measured driver of answer variance — roughly 32%, more than the model and more than the brand. English and Afrikaans phrasings of one question belong to the same prompt group and report both together and separately.
What happens when a measurement does not happen.
A vendor fails mid-cycle
The cycle completes on what it captured. Affected figures carry their actual, lower run counts — so the band widens rather than the number lying — and the report states which engines contributed.
An AI Overview is not detected
Recorded as not_detected, a third state alongside present and absent. Best-in-class vendor detection runs near 68%, so recording a miss as “absent” would permanently corrupt the trendline with false negatives.
A cycle stalls
A reconciler closes it and an alert fires. A hung cycle nobody is told about is a dashboard that looks stable because it stopped updating — the worst possible failure for a monitoring product.
The things we would put on a competitor’s page.
Published because a methodology page with no limitations section is marketing wearing a lab coat.
| Limitation | Consequence | What we do about it |
|---|---|---|
| Per-prompt estimates are directional | Three samples cannot support a per-prompt percentage | We publish a stability class instead of a number we cannot defend |
| Grounded APIs are not consumer surfaces | Shorter answers, fewer sources, search does not always fire | Instrument declared per figure; search_triggered recorded; consumer-interface calibration on the roadmap |
| AI Overview detection is imperfect | Roughly a third of present overviews are missed by vendors | Tri-state presence; undetected observations excluded from rates and counted on the report |
| The prompt set is chosen by a human | Every figure is a share of prompts someone selected | A published authoring recipe, and the full prompt list printed with the report |
| Model updates move everything at once | A step change can look like a campaign result | Engine version on every capture; version changes annotated on trendlines; optional control brands |
| Correlation, not causation | No content technique has a demonstrated stable causal effect on AI visibility under broad adoption | We measure and we say so; we do not sell the intervention |
Triangulate with data we do not control.
Every client report ships a section on checking our figures against the client’s own analytics, because a measurement you can only verify inside the tool that produced it is not verified.
- Referral traffic. AI assistants pass identifiable referrers —
utm_source=chatgpt.comand equivalents. Build the channel group; expected volumes are small, typically 0.1–0.5% of organic. - Conversion quality. AI-referred visitors have been observed converting at 4–16× the rate of general organic traffic. Small volume, disproportionate value — which is the actual commercial case for this whole discipline.
- Crawler logs. Whether the AI bots are reaching the site at all. This is the only signal in the discipline that is free of sampling noise.
- Search Console’s generative-AI report. An official anchor, and a limited one — it blends AI Overviews with AI Mode and exposes neither queries nor clicks, which is why third-party measurement remains necessary.
Bring the hardest question you have about this page
A working demo runs your prompts against live engines and we open the raw answers in front of you. If a figure does not survive that, we would rather find out with you than after the contract.