Skip to content
← Field notes
Measurement2026-08-20 · 8 min

Don't measure once: what three samples a week actually buys

The statistics say per-prompt certainty is unaffordable. They also say you do not need it.

The strongest measurement study to come out of this field in 2026 landed with an inconvenient finding: stabilising a *single prompt's* brand-detection estimate to a standard error under 0.10 takes roughly seven runs a day. Not a week. A day.

Taken at face value, that kills every commercially viable sampling design in the category, ours included. It is worth being explicit about why we did not change the design, and what we changed instead.

The arithmetic is not negotiable

Brand presence is a binomial outcome. At a true appearance rate of 30%, even a hundred runs gives you roughly ±9 percentage points. Three runs gives you almost nothing about that specific prompt. This is not a modelling choice; it is what three observations contain.

But the contracted deliverable is the aggregate

Nobody buys a report about one prompt. Fifty prompts at three samples is about 150 observations per engine per cycle, which supports a share-of-voice band of roughly ±7–8 points. That is a usable, honest number — provided it is reported as a band. Reported as a single figure, the same 150 observations become a lie with a decimal point.

Past five repeats, breadth wins

The variance-decomposition work is the finding that actually changed our design. Once you have around five samples of a phrasing, the next call reduces measurement error four to fifteen times more if you spend it on a *different phrasing or language* than on another repeat of the same one. Language alone explains roughly a third of the variance in answers — more than the model, more than the brand.

So the fix for “three is not enough” is not thirty. It is three samples across two or three phrasings, in the languages the market actually uses. That is why prompt groups and languages are first-class objects in our schema rather than a report-time convenience.

Most cells are not volatile at all

Roughly 77% of brand × prompt × engine combinations are near-deterministic — the brand is always mentioned, or never. Volatility is concentrated in a contested minority. Sampling every cell to the same depth spends most of the budget confirming things that were never in doubt, which is an argument for adaptive sampling later and for a stability class now.

What we changed

  • Every figure became a band with a run count.
  • Trend comparisons became fortnight-over-fortnight, because that is the shortest window in which movement can exceed sampling.
  • Per-prompt output became a stability class — always, contested, never — rather than a per-prompt percentage the sample cannot support.
  • Paraphrases and languages became schema, not a setting.

The sampling design did not change. The honesty of what we say about it did — and the result is a product that is *more* defensible than competitors reporting single numbers off comparable samples.

Next step

See it against your own client list

A working demo runs your prompts, in your market, on live engines — not a sandbox with seeded data. Bring one client brand and three competitors.