Built here, for the people already doing this work.
RhinocerosAI exists because the tools measuring AI search were built somewhere else, for someone else, and because the way they report is quietly indefensible. Both problems turned out to be the same problem.
Two gaps that turned out to be one.
The first gap is geographic. The tools defining this category are US- and UK-built and point a generic, English-default tracker at every market on earth. South Africa gets a two-letter country code and nothing else: no engine mix weighted to how people here actually search, no Afrikaans or isiZulu, no modelling of which AI surfaces are even available yet, and no acknowledgement that a local-intent query returns a map pack rather than an AI answer.
The second gap is methodological. The category’s sharpest criticism — three tools, the same prompts, three different answers — has never been answered. Almost every product presents a clean score and almost none will show you the observation behind it.
These look like separate problems. They are the same one: a tool that will not tell you how it measured cannot tell you it measured your market correctly either. Both are solved by the same discipline — declare the instrument, store the evidence, report what the sample supports.
What we are not
- Not a content tool. No generation, no automated fixes, no “done-for-you” tier — permanently. A vendor that scores your client and sells them the fix is a competitor with a head start.
- Not a direct-to-brand product. The agency is the customer, not a channel we tolerate until direct sales work.
- Not the first SA tool. We were, for a while, going to say we were. Then we did the research properly and found we were not. That correction is on the record.
- Not claiming to be what your customer sees. Nobody can defend that claim. We claim something narrower and checkable instead.
The decisions that keep getting made the same way.
Publish the instrument
Every figure declares how it was captured, and the methodology page is public and versioned. If we change how a number is computed, it lands in the changelog with a date, so a moved figure can always be attributed to the world or to us.
Report what the sample supports
Bands with run counts, never bare percentages. Stability classes where a percentage would overclaim. Fortnightly comparison windows, because that is the shortest honest one. This costs us demos and it is not negotiable.
Buy the hard surfaces
Grounded APIs and a search vendor, not scrapers. Lower fidelity, stated plainly, in exchange for something an agency can put in a client contract without inheriting terms-of-service exposure.
Store the evidence first
Every answer is written to storage with its provenance before anything is derived from it. It is the expensive choice — roughly 2.7 GB per brand-year — and it is what makes every other claim on this site checkable.
Publish the unflattering finding
The headline GEO paper did not replicate. llms.txt does not work. Grounded APIs are lower fidelity than consumer interfaces. Our own benchmark’s first edition is too small to be a benchmark. All of that is on this site.
A market is a build, not a toggle
We will not add a country by changing a locale string. Adding one means validating every adapter against it, establishing surface availability, writing intent rules and seeding a benchmark. One market done properly beats eighty done nominally.
Audited before it was launched, not after.
Before the platform went anywhere near a client’s data, it was put through four structured audits — security, reliability, observability, and privacy and compliance — producing 49 findings and two deploy gates that had to close before launch.
Both gates were about the same thing: whether a failure would be visible. One was the monitoring loop’s ability to run reliably at all on the deployment target. The other was that logs had nowhere durable to go, meaning an overnight failure would be gone by morning and any reliability claim would be unverifiable after the fact. Neither is a feature. Both are the difference between a monitoring product and a dashboard that has quietly stopped updating.
The stack, for the technically curious
| Concern | Choice |
|---|---|
| Application | Next.js on Vercel, TypeScript throughout |
| Database | Neon Postgres, EU region |
| Capture storage | Cloudflare R2 |
| Worker tier | Durable step functions with native fan-out |
| Auth and tenancy | Multi-tenant, agency → brand hierarchy |
| Acquisition | Grounded model APIs + SERP vendor |
The monitoring loop is a batch workload, not request/response. It runs on a dedicated worker tier for the same reason a cycle is idempotent and partial-tolerant: this is the component whose failure is invisible until someone asks why the number stopped moving.
A small team, in South Africa.
Built by people who have run agency retainers and written measurement code, which is a narrower overlap than it sounds and explains most of the product’s opinions.
Names, roles and photographs go here before launch. We have left it blank rather than filled it with placeholders — a page arguing that unverifiable claims are the problem should not open with one.
Read the arguments rather than the pitch.
See it against your own client list
A working demo runs your prompts, in your market, on live engines — not a sandbox with seeded data. Bring one client brand and three competitors.