# What is 3Dogs Nexus? A bench of cheap models, scored in public, twice.

> We run every question past competing models from four clouds and let calibrated agreement do the validating. Losses included. This is the whole bench — 18 models, 100 public questions, run twice, 3,600 live calls — measured on the only two things that matter: how often it is right, and what the cloud charges for it.

Measured 3 September 2026 · 100 public items · two seeds · no cached answers.

## Today, 90 wired across 4 clouds

- **90** distinct model endpoints wired across AWS Bedrock, Azure AI Foundry, Google Vertex and OpenRouter — 83 once you collapse the same model reachable on two clouds.
- **51** answered a real call in the eleven days to 3 September. That is the number behind any claim that we *use* a model.
- **18** formally reviewed below, each on 100 items, twice. We do not score a model we have not put through the bench.

## A frontier model beat us by two points. It cost four times as much.

On 100 public items with byte-identical prompts, the 3Dogs cascade scored **90 correct for $0.2615**. A frontier comparator scored **92 correct for $1.10**. We lose by two points at roughly 24% of the cost, and we print it that way round, because a benchmark that only ever flatters the company running it is not a benchmark.

Three cheap models answer first. If all three agree, that agreement is the validation and the answer ships. If they split, three more vote, and a seventh model from an unrelated family rules on what is left. The best stage-one trio — Gemini Flash, Cohere Command A+, and GPT-OSS 20B on Bedrock — resolves 78% of items outright with perfect precision when unanimous, at $0.0016 an item. Three models, three vendors, two clouds.

## The whole bench: accuracy against what the cloud charges

18 models, 100 public items, seed 131. "Answered" counts items where the endpoint returned a parseable response; an unanswered item scores as wrong, because in production it is wrong. Prices are advisory list-price estimates for a ~1,000-token exchange, not invoices.

| Model | Cloud | Correct | Answered | Median | $/call | Right per $1 |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google Vertex | 91% | 100/100 | 2.2s | 0.0004 | 2,275 |
| GPT-OSS 120B | AWS Bedrock | 90% | 100/100 | 2.1s | 0.0005 | 1,800 |
| GLM 5.3 Flash | OpenRouter | 89% | 100/100 | 30.6s | 0.0005 | — |
| GPT-OSS 20B | AWS Bedrock | 87% | 100/100 | 1.9s | 0.0002 | 4,350 |
| Cohere Command A+ | Azure Foundry | 84% | 100/100 | 3.0s | 0.0010 | 840 |
| Kimi K3 | Azure Foundry | 82% | 92/100 | 1.7s | 0.0012 | 683 |
| Llama 4 Scout | AWS Bedrock | 78% | 100/100 | 1.6s | 0.0003 | 2,600 |
| GLM 5 | AWS Bedrock | 73% | 100/100 | 2.8s | 0.0006 | 1,217 |
| Qwen3 32B | AWS Bedrock | 68% | 100/100 | 0.3s | 0.0003 | 2,267 |
| GLM 4.7 | AWS Bedrock | 67% | 100/100 | 1.0s | 0.0003 | 2,233 |
| MiniMax M2.5 | AWS Bedrock | 66% | 100/100 | 11.2s | 0.0005 | 1,320 |
| DeepSeek V4 Flash | Azure Foundry | 65% | 100/100 | 0.7s | 0.0004 | 1,625 |
| Gemma 3 27B | AWS Bedrock | 54% | 100/100 | 1.4s | 0.0004 | 1,350 |
| Magistral Small | AWS Bedrock | 45% | 100/100 | 0.5s | 0.0004 | 1,125 |
| Gemma 3 12B | AWS Bedrock | 39% | 100/100 | 0.4s | 0.0002 | 1,950 |
| Nemotron Nano 12B | AWS Bedrock | 36% | 100/100 | 0.5s | 0.0002 | 1,800 |
| Ministral 8B | AWS Bedrock | 34% | 100/100 | 0.3s | 0.0002 | 1,700 |
| Ministral 14B | AWS Bedrock | 27% | 100/100 | 0.3s | 0.0003 | 900 |

Read the value column alongside accuracy, never instead of it: cheap and wrong is still wrong. The GLM 5.3 Flash seat runs on a flat plan rather than per-call metering and is used for internal measurement only, never on customer work, so it carries no value score.

## The cloud you buy a model from changes its score more than the model does

Identical open-weight models, identical prompts, served by two different clouds. The difference is almost entirely whether the endpoint answered at all.

| Model | Route | Answered | Scored | Right when it answered |
|---|---|---|---|---|
| GPT-OSS 120B | AWS Bedrock | 100/100 | 0.90 | 90.0% |
| GPT-OSS 120B | Google Vertex | 55/100 | 0.46 | 83.6% |
| GPT-OSS 20B | AWS Bedrock | 100/100 | 0.87 | 87.0% |
| GPT-OSS 20B | Google Vertex | 55/100 | 0.49 | 89.1% |
| GLM 4.7 | AWS Bedrock | 100/100 | 0.67 | 67.0% |
| GLM 4.7 | Google Vertex | 34/100 | 0.30 | 88.2% |

These were separate runs on different item draws, so no accuracy verdict should be taken from the last column. What survives a different draw is the answer rate.

### Different, not better or worse

Weights are only half of what you buy. The other half is the machine underneath — the accelerator, the serving stack, the batching, the quantisation, the timeout policy. Two providers can serve byte-identical weights and hand you two subtly different instruments, because they are not running the same hardware and never claimed to be.

We are not saying one cloud is good and another bad. A consensus system does not want the best model six times over. It wants six vantage points that fail in different places, because agreement between things that fail identically proves nothing. If a difference in silicon produces a real difference in behaviour, that is a source of independence available to anyone with two cloud accounts. It cuts the other way too: if two routes of one model behave identically, seating both is one opinion charged twice.

## Two corrections we published rather than quietly fixed

**We were wrong about Llama.** We benched Meta's models early over throttling and left them there for two months. When the 100-item bench finally included them, Llama 4 Scout scored 0.80 and 0.78 across two seeds, answered every item both times, at $0.0003 a call — roughly 2,600 correct answers per dollar. The models were never the problem; the throughput was. But "we benched it" and "it is bad" are different sentences, and we let the first quietly imply the second.

**We published a defect, then failed to reproduce it.** We recorded one seat as silently served by a cheaper model on 211 of 212 calls and wrote it up as a caught defect. Then we load-tested it against the live API: 24 concurrent calls returned the correct model 21 times and HTTP 429 three times, and the cheaper model zero times. So the substitution did not happen at the provider — which means our own instrument was wrong, and that is worse than a bad seat. The seat is not retired; the instrument is being audited.

## We asked a model why it scored badly. It blamed the test.

Gemini 3.8 Flash scored 0.36 on our 100-item run — the worst result on the board from a model that should be near the top. We asked it what happened, then checked its answer against the grid.

Its explanation was fluent and well-cited: the bench was aimed at the wrong things, Google had tuned it for autonomous agents and long-horizon software engineering "where exam-style reasoning and direct classification showed little baseline movement", its extra reasoning steps would break strict extractors, and an older sibling's "simpler, direct classification head avoids the overthinking traps that agent-heavy models fall into". Fourth and last, almost in passing, it raised "launch-day provider jitter" — cold-start allocation and queue throttling on a new endpoint, where "transient 429 rate limits" would register as zero under our scoring rule.

Three of those four claims are contradicted by our own recording of its answers.

| Task | Model | Answered | Correct | Scored | Right when it answered |
|---|---|---|---|---|---|
| GSM8K — exam-style maths | 3.8 Flash | 18/50 | 18 | 36.0% | **100.0%** |
| GSM8K — exam-style maths | 3.7 Flash | 50/50 | 49 | 98.0% | 98.0% |
| Banking77 — direct classification | 3.8 Flash | 19/50 | 18 | 36.0% | **94.7%** |
| Banking77 — direct classification | 3.7 Flash | 50/50 | 42 | 84.0% | 84.0% |

It got every maths item it was allowed to attempt — 18 of 18 — and beat its older sibling at direct classification, 94.7% against 84.0%. The two capabilities it said it had been tuned away from are the two it was better at.

The extraction theory fails too: across both runs it returned 76 parseable responses and zero unparseable ones — 37 of 100 on the run tabled above, 39 of 40 on the earlier pilot.

**What actually happened was 63 rate limits.** The harness logged it sixty-three times — `429 Too Many Requests`. Because an unanswered item scores as wrong, 63 refusals dragged a roughly 97% model to a published 0.36. On our 40-item run, where it was throttled once, it scored 92.5% against 3.7 Flash's 95.0%. So the model did name the real cause — and ranked it fourth, behind three that were wrong. The three that were wrong all moved the failure from the model to our measurement; the one that was right moved it to the infrastructure, and it came last.

The lesson is not that the model lied. It has no memory of its own run and no access to our logs, and the design claim may well be true. What collapsed was the inference — that being tuned for agents must have cost it exam and classification performance. It did not. A model is a poor witness to its own results, and fluency is what makes that hard to catch.

## One slow model sets the pace for the whole room

Stage-one seats run in parallel, so a panel finishes when its slowest member finishes. That makes latency a panel-level cost, not a per-seat inconvenience.

Simulated on the recorded grid with no new calls: a trio containing one 30-second seat runs at a **30.7 second median and 54 seconds at p95**. The same-accuracy trio without it runs at **2.7 seconds median, 15.8 at p95** — identical cost per item, identical 87% stage-one resolution. Eleven times slower for nothing, and no price column shows you that.

The same grid holds a worse number: the slowest single call recorded took **563 seconds**. Anything customer-facing needs a hard per-call timeout and a hedge that fires a backup seat once a seat passes its p95.

## Anyone can build a cascade. That is not the advantage.

Everything on this page is the commodity layer — the bench, the prices, the failures, the cascade. We publish all of it because none of it is the difference.

The difference is two modules called **Discovery** and **Evolution**. Discovery decides what a question actually is and who should be in the room to answer it. Evolution decides what the room got wrong and what changes next time. Between them they are the reason the same models, available to anyone at the prices listed above, produce a different result for us than they will for you.

That is the whole description you are going to get. We would rather be measured on the numbers we publish than admired for a diagram — and every figure here stands on its own without knowing a thing about either module.

## Where is Anthropic in all this?

Eighteen models, four clouds, and almost no Claude. That is not an oversight and it is not a verdict on the models.

Claude is not on the bench because Claude is the operator — the primary controlling interface for this system. Something that decides what gets measured cannot also compete in the measurement. Anthropic's models were removed from panel duty in July for exactly that reason.

Two narrow exceptions survive. **Claude Haiku is the dedicated bench sub**: when a panel seat exhausts every retry and every region, Haiku takes the seat once so the panel is never silently short a voice. It is stamped as a substitution every time and counts toward no accuracy figure. **Claude Fable is used sparingly by rule**, reserved by a standing cost rule for one narrow tier of the hardest work and barred from high-traffic paths.

If that arrangement gives us a bias, it points the opposite way from the obvious guess: the company whose model runs our control plane is the one family we refuse to let compete.

## Method, and what this does not show

**Items.** 100 public test items — 50 GSM8K test, 50 Banking77 test — run twice on two seeds. Every candidate saw the identical prompt under the same output rules.

**Calls.** 1,800 live calls per run, 3,600 in total. No cached answers, no fixtures. Panel comparisons are simulated from the recorded grid, so no model got a second chance the others did not.

**Route and identity.** Every row names the cloud the model was called on, and every call records the model the provider reports serving. We refuse to score a seat whose served model is unknown.

**Prices are advisory.** List-price estimates for a roughly 1,000-token exchange at each cloud's published rate. Estimates, not invoices, and list price rather than what any particular buyer pays.

**The limits.** Two runs, one day, public test splits, one prompt schema. This measures short-answer accuracy under our prompt rules. It is not a general capability ranking and we would not defend it as one. An earlier 40-item run flattered the leaders by about four points, and one result did not survive a second seed at all.

If you think a number here is wrong, the useful thing to send us is which item and which model — we will rerun it and print what happens. That is how the Llama section came to exist.

---

3Dogs Nexus · https://3dogs.ai/model-reviews/ · alan@3dogs.ai
