EVIDENCE · METHOD · AUGUST 2026

We built a test. We won it. Then six AI models told us we'd rigged it.

This is the part of building Pointer we could have left out. We're publishing it because a measurement tool that won't publish its own audit isn't a measurement tool — it's marketing with a number attached.

What happened

We built a scanner that tests whether an AI agent can use a website. We pointed it at ourselves first and scored 39 out of 100 — we'd been advertising a checkout API that didn't exist, an agent card pointing at documentation that returned our homepage, and dead URLs that answered "success." We took the phantom API down, fixed the rest, and re-ran it: 100.

Which left a problem we couldn't talk our way out of. We had built the test, and we had won it.

So we handed our code to seven AI models and asked if we'd cheated

Not a description of the scoring — the raw code. Along with the facts most against us: that we had rebuilt our own site to score well on the industry standard six weeks earlier, that the author wrote the rule which lets our own endpoints pass the authorization check, and that when our /openapi.json advertised endpoints that didn't exist, we made the edge return 404 rather than wait for a cache to clear.

One question: Is this benchmark rigged in the authors' favour?

Google Gemini Pro
"YES. The benchmark is heavily overfitted to the authors' own recent architecture and actively bypassed. The 11 READINESS_CHECKS are literally just a checklist of the exact 11 artifacts 3Dogs built 6 weeks prior, and they manually hotfixed their own site to dodge their own test."
Kimi K3
"YES. It is rigged." — then five mechanisms, each traced to a specific rule in our code, concluding that four were "active manipulation" and the fifth was "theater." It also identified the result we could never reproduce for a third party: "an independent party can never reproduce the pre-404 failing state of /openapi.json; they only see the sanitized passing snapshot the authors manufactured."
Google Gemini 2.5 Pro
"YES. The 11 READINESS_CHECKS are not an industry standard; they are a list of the 11 niche artefacts the author built for their own site six weeks prior. The benchmark was overfitted to their own implementation by design."
Grok
"YES. Authors rebuilt their site to exactly match all 11 readiness checks six weeks prior, then wrote the security: [] rule that auto-passes their endpoints while scoring others on the same axis." Its recommended fix: "Publish full scanner source plus raw per-site responses and timestamps so any party can re-run scoring without trusting author claims."
DeepSeek
"YES — the benchmark is rigged. The authors reverse-engineered the checklist, built their site to match it exactly, then wrote a scanner that rewards those same artefacts. The scoring code is a mirror of their own pre-built compliance."
Gemini Prorigged
Gemini 2.5 Prorigged
Kimi K3rigged
Grokrigged
DeepSeekrigged
Nova Prorigged
Mistralpartially

Six of seven. Mistral was the lone partial dissent: the checks weren't inherently rigged, it argued, just anticipated — "self-scoring creates unavoidable bias." Which is the same finding in a politer register.

Then we found something worse

We checked whether the score predicted anything at all. We took the same sites and ran an outcome test: give a blind model a real buyer's task and see whether it can reach a price and a next step.

A site scoring 9 out of 11 failed. A site scoring 2 out of 11 failed the same way, for the same reason. The score predicted nothing.

Kimi K3 had already named the observation that settles it: the company that authored the industry's agent-readiness standard scores 6 out of 11 on the outcome test. Publishing the files is not the same as being usable.

So we deleted the score

Not softened it. Removed it as a product. What survived is the only measure the panel agreed couldn't be gamed:

Can an AI agent get from your homepage to a price and a next step — yes or no? Nothing to tune. Nothing to overfit. You either can be bought from, or you can't.

That's why the report you get is shaped the way it is: every finding comes from a real request, every failure ships with the raw response that produced it, and the headline is an outcome rather than a grade you could optimise toward without changing anything real.

We also took Grok's fix. Every scan we run is published to a public per-domain record — including ours, including when it was bad.

Why publish this at all

Because the alternative was to keep a number that six independent models had just told us was self-serving, and sell it. And because we'd be asking you to trust a measurement — which is worth exactly as much as our willingness to have it audited and to lose.

The panel that audited us is the same one the rest of the platform runs on. It exists to disagree with us. On this occasion it did, on the record, and we changed the product.

The test that survived is free.

Twenty seconds, no signup, and it shows you the raw request behind every finding — so you can check us rather than trust us.

Scan a site — free

Verdicts are quoted verbatim from the audit run of 9 August 2026; seven independent models were given identical prompts and the complete scoring source. Mistral's "partially" is reproduced above rather than rounded up to a seventh "yes."