What is 3Dogs Nexus? A bench of cheap models, scored in public, twice.
We run every question past competing models from four clouds and let calibrated agreement do the validating. Losses included. This is the whole bench — 18 models, 100 public questions, run twice, 3,600 live calls — measured on the only two things that matter: how often it is right, and what the cloud charges for it.
Measured 3 September 2026 · 100 public items · two seeds · no cached answers · machine-readable version
Today, 90 wired across 4 clouds
“How many models do you use” deserves an arithmetic answer rather than a marketing one. Here is ours, and exactly what each number means.
A frontier model beat us by two points. It cost four times as much.
On 100 public items with byte-identical prompts, our cascade scored 90 correct for $0.2615. A frontier comparator scored 92 correct for $1.10. We lose by two points at roughly 24% of the cost, and we print it that way round, because a benchmark that only ever flatters the company running it is not a benchmark.
Three cheap models answer first. If all three agree, that agreement is the validation and the answer ships. If they split, three more vote, and a seventh model from an unrelated family rules on what is left. The best stage-one trio — Gemini Flash, Cohere Command A+, and GPT-OSS 20B on Bedrock — resolves 78% of items outright with perfect precision when unanimous, at $0.0016 an item. Three models, three vendors, two clouds.
Accuracy against what the cloud charges
18 models, 100 public items, seed 131. “Answered” counts items where the endpoint returned a parseable response; an unanswered item scores as wrong, because in production it is wrong.
| Model | Cloud | Correct | Answered | Median | $/call | Right per $1 |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | Google Vertex | 91% | 100/100 | 2.2s | 0.0004 | 2,275 |
| GPT-OSS 120B | AWS Bedrock | 90% | 100/100 | 2.1s | 0.0005 | 1,800 |
| GLM 5.3 Flash | OpenRouter | 89% | 100/100 | 30.6s | 0.0005 | — |
| GPT-OSS 20B | AWS Bedrock | 87% | 100/100 | 1.9s | 0.0002 | 4,350 |
| Cohere Command A+ | Azure Foundry | 84% | 100/100 | 3.0s | 0.0010 | 840 |
| Kimi K3 | Azure Foundry | 82% | 92/100 | 1.7s | 0.0012 | 683 |
| Llama 4 Scout | AWS Bedrock | 78% | 100/100 | 1.6s | 0.0003 | 2,600 |
| GLM 5 | AWS Bedrock | 73% | 100/100 | 2.8s | 0.0006 | 1,217 |
| Qwen3 32B | AWS Bedrock | 68% | 100/100 | 0.3s | 0.0003 | 2,267 |
| GLM 4.7 | AWS Bedrock | 67% | 100/100 | 1.0s | 0.0003 | 2,233 |
| MiniMax M2.5 | AWS Bedrock | 66% | 100/100 | 11.2s | 0.0005 | 1,320 |
| DeepSeek V4 Flash | Azure Foundry | 65% | 100/100 | 0.7s | 0.0004 | 1,625 |
| Gemma 3 27B | AWS Bedrock | 54% | 100/100 | 1.4s | 0.0004 | 1,350 |
| Magistral Small | AWS Bedrock | 45% | 100/100 | 0.5s | 0.0004 | 1,125 |
| Gemma 3 12B | AWS Bedrock | 39% | 100/100 | 0.4s | 0.0002 | 1,950 |
| Nemotron Nano 12B | AWS Bedrock | 36% | 100/100 | 0.5s | 0.0002 | 1,800 |
| Ministral 8B | AWS Bedrock | 34% | 100/100 | 0.3s | 0.0002 | 1,700 |
| Ministral 14B | AWS Bedrock | 27% | 100/100 | 0.3s | 0.0003 | 900 |
Read the value column alongside accuracy, never instead of it: cheap and wrong is still wrong. The GLM 5.3 Flash seat runs on a flat plan rather than per-call metering and is used for internal measurement only, never on customer work, so it carries no value score here.
The cloud you buy a model from changes its score more than the model does
Identical open-weight models, identical prompts, served by two different clouds. The difference is almost entirely whether the endpoint answered at all.
| Model | Route | Answered | Scored | Right when it answered |
|---|---|---|---|---|
| GPT-OSS 120B | AWS Bedrock | 100/100 | 0.90 | 90.0% |
| GPT-OSS 120B | Google Vertex | 55/100 | 0.46 | 83.6% |
| GPT-OSS 20B | AWS Bedrock | 100/100 | 0.87 | 87.0% |
| GPT-OSS 20B | Google Vertex | 55/100 | 0.49 | 89.1% |
| GLM 4.7 | AWS Bedrock | 100/100 | 0.67 | 67.0% |
| GLM 4.7 | Google Vertex | 34/100 | 0.30 | 88.2% |
Different, not better or worse
Weights are only half of what you buy. The other half is the machine underneath — the accelerator, the serving stack, the batching, the quantisation, the timeout policy. Two providers can serve byte-identical weights and hand you two subtly different instruments, because they are not running the same hardware and never claimed to be.
We are not saying one cloud is good and another bad. We are saying they are not the same thing, and for us that is the useful part. A consensus system does not want the best model six times over. It wants six vantage points that fail in different places, because agreement between things that fail identically proves nothing. If a difference in silicon produces a real difference in behaviour, that is a source of independence available to anyone with two cloud accounts.
It cuts the other way too, which is why we are careful: if two routes of one model behave identically, seating both is one opinion charged twice. Whether they genuinely disagree is a measurement we have not made yet, and we would rather say so than imply we have.
Two corrections we published rather than quietly fixed
We were wrong about Llama
We benched Meta's models early over throttling and left them there for two months. When the 100-item bench finally included them, Llama 4 Scout scored 0.80 and 0.78 across two seeds, answered every single item both times, at $0.0003 a call — roughly 2,600 correct answers per dollar, better value than the most accurate seat on the board.
The models were never the problem; the throughput was. But “we benched it” and “it is bad” are different sentences, and we let the first quietly imply the second.
We published a defect, then failed to reproduce it
We recorded one seat as silently served by a cheaper model on 211 of 212 calls and wrote it up as a caught defect. Then we load-tested it against the live API: 24 concurrent calls returned the correct model 21 times and HTTP 429 three times, and the cheaper model zero times.
So the substitution did not happen at the provider — which means our own instrument was wrong, and that is worse than a bad seat. The seat is not retired; the instrument is being audited. We leave this on the page because a benchmark that shows only its clean findings is telling you how it edits, not how it measures.
We asked a model why it scored badly. It blamed the test.
Gemini 3.8 Flash scored 0.36 on our 100-item run — the worst result on the board from a model that should be near the top. So we asked it what happened, then checked its answer against the grid.
Its explanation was fluent, specific and well-cited. The bench was aimed at the wrong things, it said: Google had tuned it for autonomous agents and long-horizon software engineering, “where exam-style reasoning and direct classification showed little baseline movement”; its extra internal reasoning steps would break strict output extractors; and an older sibling’s “simpler, direct classification head avoids the overthinking traps that agent-heavy models fall into.” Fourth and last, almost in passing, it raised “launch-day provider jitter” — cold-start allocation and queue throttling on a new endpoint, where “transient 429 rate limits” would register as zero under our scoring rule.
Three of those four claims are contradicted by our own recording of its answers.
| Task | Model | Answered | Correct | Scored | Right when it answered |
|---|---|---|---|---|---|
| GSM8K — exam-style maths | 3.8 Flash | 18/50 | 18 | 36.0% | 100.0% |
| GSM8K — exam-style maths | 3.7 Flash | 50/50 | 49 | 98.0% | 98.0% |
| Banking77 — direct classification | 3.8 Flash | 19/50 | 18 | 36.0% | 94.7% |
| Banking77 — direct classification | 3.7 Flash | 50/50 | 42 | 84.0% | 84.0% |
It got every maths item it was allowed to attempt — 18 of 18 — and beat its older sibling at direct classification too, 94.7% against 84.0%. The two capabilities it told us it had been tuned away from are the two it was better at, and the “overthinking trap” it attributed to itself appears in none of the 37 answers it returned.
The extraction theory fails as cleanly: across both runs it returned 76 parseable responses and zero unparseable ones — 37 of 100 on the run tabled above, 39 of 40 on the earlier pilot. If reasoning overhead were breaking our harness we would see answers that arrived and could not be scored. There are none.
What actually happened was 63 rate limits. The harness logged the reason itself, sixty-three times: 429 Too Many Requests. The model was throttled on a freshly launched endpoint, and because an unanswered item scores as wrong, 63 refusals dragged a roughly 97% model down to a published 0.36. On our earlier 40-item run, where it was throttled exactly once, it scored 92.5% against 3.7 Flash’s 95.0% — a gap nobody would write an essay about. So the model did name the real cause — and ranked it fourth, behind three that were wrong. The three that were wrong had something in common: each moved the failure from the model to our measurement. The one that was right moved it to the infrastructure, and it came last.
The lesson is not that the model lied. It has no memory of its own run and no access to our logs, and the design claim may well be true — Google may genuinely have tuned it for agent work. What collapsed was the inference: that being tuned for agents must have cost it exam and classification performance. It did not. A model is a poor witness to its own results, and the fluency of the account is exactly what makes it hard to catch. That is why every number on this page carries the route it came from, and why we refuse to score a seat whose served model we cannot confirm.
One slow model sets the pace for the whole room
Stage-one seats run in parallel, so a panel finishes when its slowest member finishes. That makes latency a panel-level cost, not a per-seat inconvenience.
Simulated on the recorded grid with no new calls: a trio containing one 30-second seat runs at a 30.7 second median and 54 seconds at p95. The same-accuracy trio without it runs at 2.7 seconds median, 15.8 at p95 — identical cost per item, identical 87% stage-one resolution. Eleven times slower for nothing, and no price column shows you that.
The same grid holds a worse number: the slowest single call we recorded took 563 seconds. One seat hung for over nine minutes. Medians are irrelevant to whoever is waiting, so anything customer-facing needs a hard per-call timeout and a hedge that fires a backup seat once a seat passes its p95. We treat that as a precondition, not an improvement.
Anyone can build a cascade. That is not the advantage.
Everything on this page is the commodity layer — the bench, the prices, the failures, the cascade. We publish all of it because none of it is the difference.
The difference is two modules called Discovery and Evolution. Discovery decides what a question actually is and who should be in the room to answer it. Evolution decides what the room got wrong and what changes next time. Between them they are the reason the same models, available to anyone at the prices listed above, produce a different result for us than they will for you.
That is the whole description you are going to get. Not the architecture, not the seat bindings, not the taxonomy underneath either, and not the loop that connects them. We would rather be measured on the numbers we publish than admired for a diagram — and every figure on this page stands on its own without knowing a thing about either module.
Where is Anthropic in all this?
Eighteen models, four clouds, and almost no Claude. That is not an oversight and it is not a verdict on the models.
Claude is not on the bench because Claude is the operator. It is the primary controlling interface for this system — it reads the state, writes the specs, dispatches the work, and reviews what comes back. Something that decides what gets measured cannot also compete in the measurement. We removed Anthropic's models from panel duty in July for exactly that reason, and the scoreboard has stayed clear of the scorer since.
Two narrow exceptions survive. Claude Haiku is the dedicated bench sub: when a panel seat exhausts every retry and every region, Haiku comes off the bench and takes the seat once, so the panel is never silently short a voice. It is stamped as a substitution every time it plays and counts toward no accuracy figure here — a substitute's minutes do not go on the starter's average. And Claude Fable is used sparingly by rule, reserved by a standing cost rule for one narrow tier of the hardest work and barred from the high-traffic paths.
If that arrangement gives us a bias, it points the opposite way from the obvious guess: the company whose model runs our control plane is the one family we refuse to let compete.
How this was measured, and what it does not show
What we did
Items. 100 public test items — 50 GSM8K test, 50 Banking77 test — run twice on two seeds. Every candidate saw the identical prompt under the same output rules.
Calls. 1,800 live calls per run, 3,600 in total. No cached answers, no fixtures. Panel comparisons are simulated from the recorded grid, so no model got a second chance the others did not.
Route and identity. Every row names the cloud the model was called on, and every call records the model the provider reports serving. We refuse to score a seat whose served model is unknown.
What it does not show
Prices are advisory. List-price estimates for a roughly 1,000-token exchange at each cloud's published rate. Estimates, not invoices, and list price rather than what any particular buyer pays.
The limits. Two runs, one day, public test splits, one prompt schema. This measures short-answer accuracy under our prompt rules. It is not a general capability ranking and we would not defend it as one.
Replication moved the numbers. An earlier 40-item run flattered the leaders by about four points, and one result did not survive a second seed at all. Both are printed rather than dropped.
If you think a number here is wrong, the useful thing to send us is which item and which model — we will rerun it and print what happens. That is how the Llama section above came to exist.