Skip to content
Edition 0 · commissioning data · 2026-08-27

The SA AI Search Benchmark

There is no credible published measurement of how AI search actually behaves in South Africa — no local AI Overview trigger rates, no engine-agreement data, no language effects. We generate thousands of South African observations a week as a by-product of doing the job, so we publish the aggregate. This is the first edition, and it is small. We are going to say exactly how small.

00Read this first

Edition 0 is not yet a benchmark.

It is the commissioning dataset from bringing the platform live: hundreds of runs rather than tens of thousands, concentrated on one South African brand category and one engine leg. Every figure below carries its actual sample size, and several of them are too small to support a market-level claim. They are published anyway, because a benchmark that starts by pretending to be finished is exactly the behaviour this platform exists to argue against.

Edition 1 arrives once the weekly cycle has run across a multi-brand cohort on all six engines. What that edition will contain is listed at the bottom of this page, so you can hold us to it.

01The instrument
Period2026-W34 to W35
MarketZA · en
Engine legsClaude (grounded), live batch
Runs210 measurement + 18 verification
Cited sources618 observed
Samplesn=3 per prompt

A single-engine dataset cannot speak to engine agreement, and a single-category dataset cannot speak to a national trigger rate. Both gaps are named where they matter below.

02Finding 01

Owned-domain sets are systematically understated.

The strongest finding in the commissioning data is not about AI engines at all. It is about how brands are configured in tools like ours.

Across 618 cited sources for one South African brand, configuring the brand with a single owned domain — the normal default — attributed 6 citations to the brand. Configuring its actual owned-domain set attributed 31. The default understated owned-citation share by 5.2×, silently, with no error and no warning.

If you are running any AI-visibility tool today, this is the single highest-value thing to go and check. It costs ten minutes and it may be the reason a client’s owned-citation number looks hopeless.

Owned-citation share, one brand

Configured with one domain0.4–2.1%
6 of 618 cited sources
Configured with the full owned set3.5–7.1%
31 of 618 cited sources

Both bands are Wilson intervals on the same 618 observations. Note that even the corrected figure is low — which is normal. Across the field, roughly 84% of AI citations are third-party.

03Finding 02

A cheaper instrument returned identical detection.

Two acquisition configurations were run against the same prompts in the same window: a large model called synchronously, and a batch submission with a mid-tier model doing the analysis pass.

ConfigurationCost per runMention rateRuns
Large model, synchronous$0.158630/30105
Batch + mid-tier analysis$0.033730/30105
Live verification cycle$0.029318/18 completed18

What it means for buyers

Cost per run in this category is largely an engineering choice rather than a floor. A vendor charging per engine because “the models are expensive” is describing an architecture, not a law of nature.

One operational surprise

Batch latency is not proportional to batch size. Eighteen requests took 12.8 minutes; 150 took 5.0. Small batches are not fast batches — which matters if a vendor promises you on-demand refreshes.

04Finding 03

Source churn is the default state.

Two capture windows 49 minutes apart on the same prompts produced substantially different citation sets — consistent with the published finding that day-to-day overlap of cited sources for identical prompts runs at 34–42%.

The practical consequence for anyone reading a citation table: a domain that appears once is noise. A domain that appears across a fortnight is a source. Any tool presenting a single day’s citation list as “the sources shaping your category” is presenting a snapshot of a shuffle.

Storage, for anyone sizing this

MeasureObserved
Mean grounded capture size56.7 KB
Captures per brand-year (50 prompts, 6 engines, n=3, weekly)≈ 47,000
Storage per brand-year≈ 2.7 GB

Published because it is the number that decides whether a vendor can afford to keep raw answers at all — and a vendor that cannot keep them cannot let you audit anything.

05What we cannot yet say

The gaps, named.

Everything below is a question this benchmark exists to answer and cannot answer yet. Listing them is the point: it is what makes Edition 1 checkable against Edition 0.

QuestionWhy not yetEdition
South African AI Overview trigger rate by intentRequires the Google surface legs across a multi-category prompt cohort1
Do the six engines agree on who to recommend?Single-engine dataset — engine agreement is undefined on one leg1
English versus Afrikaans mention ratesAfrikaans prompt sets are authored but not yet cycled at volume1
How often grounded APIs actually search, on ZA promptssearch_triggered is recorded; the sample is not yet large enough to publish a rate1
Category-level share-of-voice concentrationNeeds a cohort of brands per category, not one brand2
Model-event impact on South African answersRequires history spanning a provider model update2
  • Edition 1 — once the weekly cycle has run across a multi-brand cohort on all six engines. Trigger rates by intent, engine agreement, and the first English/Afrikaans comparison at volume.
  • Edition 2 — category concentration, model-event effects, and the first genuine quarter-over-quarter trend.
  • Method, published with each edition. Prompt cohort, run counts, coverage, and the raw aggregates. If a figure here is ever quoted back at us, it should be checkable.
06Why we publish it

The absence is the opportunity.

South Africa has around 96% Google share, roughly 70% adult chatbot adoption, AI Mode live since August 2025 and four local languages in Google’s AI surfaces since March 2026 — and no published measurement of any of it. Every agency in this market is currently advising clients on AI search using data from somewhere else.

We are going to be running these cycles regardless. Publishing the aggregate costs us almost nothing and is the one piece of evidence a competitor cannot copy without doing the same work for a year.

If you want to contribute

Agencies running brands on the platform can opt a brand into the benchmark cohort. Contributed data is aggregated to category level and never published brand-identifiably — the benchmark reports what a category looks like, never what your client looks like.

Join the cohort
Next step

Measure your own category before the next edition

Edition 1 will describe categories in aggregate. A working demo describes yours specifically, this week, with the raw answers open in front of you.