The model bench

What is 3Dogs Nexus? A bench of cheap models, scored in public, twice.

We run every question past competing models from four clouds and let calibrated agreement do the validating. Losses included. This is the whole bench — 18 models, 100 public questions, run twice, 3,600 live calls — measured on the only two things that matter: how often it is right, and what the cloud charges for it.

Measured 3 September 2026 · 100 public items · two seeds · no cached answers · machine-readable version

The count, and the rule we count by

Today, 90 wired across 4 clouds

“How many models do you use” deserves an arithmetic answer rather than a marketing one. Here is ours, and exactly what each number means.

90distinct model endpoints wired across AWS Bedrock, Azure AI Foundry, Google Vertex and OpenRouter — 83 once you collapse the same model reachable on two clouds
51answered a real call in the eleven days to 3 September. That is the number behind any claim that we use a model
18formally reviewed below, each on 100 items, twice. We do not score a model we have not put through the bench
The result we lead with

A frontier model beat us by two points. It cost four times as much.

On 100 public items with byte-identical prompts, our cascade scored 90 correct for $0.2615. A frontier comparator scored 92 correct for $1.10. We lose by two points at roughly 24% of the cost, and we print it that way round, because a benchmark that only ever flatters the company running it is not a benchmark.

Three cheap models answer first. If all three agree, that agreement is the validation and the answer ships. If they split, three more vote, and a seventh model from an unrelated family rules on what is left. The best stage-one trio — Gemini Flash, Cohere Command A+, and GPT-OSS 20B on Bedrock — resolves 78% of items outright with perfect precision when unanimous, at $0.0016 an item. Three models, three vendors, two clouds.

The whole bench

Accuracy against what the cloud charges

18 models, 100 public items, seed 131. “Answered” counts items where the endpoint returned a parseable response; an unanswered item scores as wrong, because in production it is wrong.

3Dogs model bench, week 2026-W36. Prices are advisory list-price estimates for a ~1,000-token exchange, not invoices.
ModelCloudCorrectAnsweredMedian$/callRight per $1
Gemini 3.7 FlashGoogle Vertex91%100/1002.2s0.00042,275
GPT-OSS 120BAWS Bedrock90%100/1002.1s0.00051,800
GLM 5.3 FlashOpenRouter89%100/10030.6s0.0005
GPT-OSS 20BAWS Bedrock87%100/1001.9s0.00024,350
Cohere Command A+Azure Foundry84%100/1003.0s0.0010840
Kimi K3Azure Foundry82%92/1001.7s0.0012683
Llama 4 ScoutAWS Bedrock78%100/1001.6s0.00032,600
GLM 5AWS Bedrock73%100/1002.8s0.00061,217
Qwen3 32BAWS Bedrock68%100/1000.3s0.00032,267
GLM 4.7AWS Bedrock67%100/1001.0s0.00032,233
MiniMax M2.5AWS Bedrock66%100/10011.2s0.00051,320
DeepSeek V4 FlashAzure Foundry65%100/1000.7s0.00041,625
Gemma 3 27BAWS Bedrock54%100/1001.4s0.00041,350
Magistral SmallAWS Bedrock45%100/1000.5s0.00041,125
Gemma 3 12BAWS Bedrock39%100/1000.4s0.00021,950
Nemotron Nano 12BAWS Bedrock36%100/1000.5s0.00021,800
Ministral 8BAWS Bedrock34%100/1000.3s0.00021,700
Ministral 14BAWS Bedrock27%100/1000.3s0.0003900

Read the value column alongside accuracy, never instead of it: cheap and wrong is still wrong. The GLM 5.3 Flash seat runs on a flat plan rather than per-call metering and is used for internal measurement only, never on customer work, so it carries no value score here.

The finding nobody benchmarks

The cloud you buy a model from changes its score more than the model does

Identical open-weight models, identical prompts, served by two different clouds. The difference is almost entirely whether the endpoint answered at all.

Same weights, two clouds. Note: these were separate runs on different item draws, so no accuracy verdict should be taken from the last column.
ModelRouteAnsweredScoredRight when it answered
GPT-OSS 120BAWS Bedrock100/1000.9090.0%
GPT-OSS 120BGoogle Vertex55/1000.4683.6%
GPT-OSS 20BAWS Bedrock100/1000.8787.0%
GPT-OSS 20BGoogle Vertex55/1000.4989.1%
GLM 4.7AWS Bedrock100/1000.6767.0%
GLM 4.7Google Vertex34/1000.3088.2%

Different, not better or worse

Weights are only half of what you buy. The other half is the machine underneath — the accelerator, the serving stack, the batching, the quantisation, the timeout policy. Two providers can serve byte-identical weights and hand you two subtly different instruments, because they are not running the same hardware and never claimed to be.

We are not saying one cloud is good and another bad. We are saying they are not the same thing, and for us that is the useful part. A consensus system does not want the best model six times over. It wants six vantage points that fail in different places, because agreement between things that fail identically proves nothing. If a difference in silicon produces a real difference in behaviour, that is a source of independence available to anyone with two cloud accounts.

It cuts the other way too, which is why we are careful: if two routes of one model behave identically, seating both is one opinion charged twice. Whether they genuinely disagree is a measurement we have not made yet, and we would rather say so than imply we have.

Where we were wrong

Two corrections we published rather than quietly fixed

We were wrong about Llama

We benched Meta's models early over throttling and left them there for two months. When the 100-item bench finally included them, Llama 4 Scout scored 0.80 and 0.78 across two seeds, answered every single item both times, at $0.0003 a call — roughly 2,600 correct answers per dollar, better value than the most accurate seat on the board.

The models were never the problem; the throughput was. But “we benched it” and “it is bad” are different sentences, and we let the first quietly imply the second.

We published a defect, then failed to reproduce it

We recorded one seat as silently served by a cheaper model on 211 of 212 calls and wrote it up as a caught defect. Then we load-tested it against the live API: 24 concurrent calls returned the correct model 21 times and HTTP 429 three times, and the cheaper model zero times.

So the substitution did not happen at the provider — which means our own instrument was wrong, and that is worse than a bad seat. The seat is not retired; the instrument is being audited. We leave this on the page because a benchmark that shows only its clean findings is telling you how it edits, not how it measures.

A third correction, and the strangest

We asked a model why it scored badly. It blamed the test.

Gemini 3.8 Flash scored 0.36 on our 100-item run — the worst result on the board from a model that should be near the top. So we asked it what happened, then checked its answer against the grid.

Its explanation was fluent, specific and well-cited. The bench was aimed at the wrong things, it said: Google had tuned it for autonomous agents and long-horizon software engineering, “where exam-style reasoning and direct classification showed little baseline movement”; its extra internal reasoning steps would break strict output extractors; and an older sibling’s “simpler, direct classification head avoids the overthinking traps that agent-heavy models fall into.” Fourth and last, almost in passing, it raised “launch-day provider jitter” — cold-start allocation and queue throttling on a new endpoint, where “transient 429 rate limits” would register as zero under our scoring rule.

Three of those four claims are contradicted by our own recording of its answers.

Same 100 items, same run, same prompts — Gemini 3.8 Flash against 3.7 Flash
TaskModelAnsweredCorrectScoredRight when it answered
GSM8K — exam-style maths3.8 Flash18/501836.0%100.0%
GSM8K — exam-style maths3.7 Flash50/504998.0%98.0%
Banking77 — direct classification3.8 Flash19/501836.0%94.7%
Banking77 — direct classification3.7 Flash50/504284.0%84.0%

It got every maths item it was allowed to attempt — 18 of 18 — and beat its older sibling at direct classification too, 94.7% against 84.0%. The two capabilities it told us it had been tuned away from are the two it was better at, and the “overthinking trap” it attributed to itself appears in none of the 37 answers it returned.

The extraction theory fails as cleanly: across both runs it returned 76 parseable responses and zero unparseable ones — 37 of 100 on the run tabled above, 39 of 40 on the earlier pilot. If reasoning overhead were breaking our harness we would see answers that arrived and could not be scored. There are none.

What actually happened was 63 rate limits. The harness logged the reason itself, sixty-three times: 429 Too Many Requests. The model was throttled on a freshly launched endpoint, and because an unanswered item scores as wrong, 63 refusals dragged a roughly 97% model down to a published 0.36. On our earlier 40-item run, where it was throttled exactly once, it scored 92.5% against 3.7 Flash’s 95.0% — a gap nobody would write an essay about. So the model did name the real cause — and ranked it fourth, behind three that were wrong. The three that were wrong had something in common: each moved the failure from the model to our measurement. The one that was right moved it to the infrastructure, and it came last.

The lesson is not that the model lied. It has no memory of its own run and no access to our logs, and the design claim may well be true — Google may genuinely have tuned it for agent work. What collapsed was the inference: that being tuned for agents must have cost it exam and classification performance. It did not. A model is a poor witness to its own results, and the fluency of the account is exactly what makes it hard to catch. That is why every number on this page carries the route it came from, and why we refuse to score a seat whose served model we cannot confirm.

The cost nobody puts in the price column

One slow model sets the pace for the whole room

Stage-one seats run in parallel, so a panel finishes when its slowest member finishes. That makes latency a panel-level cost, not a per-seat inconvenience.

Simulated on the recorded grid with no new calls: a trio containing one 30-second seat runs at a 30.7 second median and 54 seconds at p95. The same-accuracy trio without it runs at 2.7 seconds median, 15.8 at p95 — identical cost per item, identical 87% stage-one resolution. Eleven times slower for nothing, and no price column shows you that.

The same grid holds a worse number: the slowest single call we recorded took 563 seconds. One seat hung for over nine minutes. Medians are irrelevant to whoever is waiting, so anything customer-facing needs a hard per-call timeout and a hedge that fires a backup seat once a seat passes its p95. We treat that as a precondition, not an improvement.

Two modules we will not describe

Anyone can build a cascade. That is not the advantage.

Everything on this page is the commodity layer — the bench, the prices, the failures, the cascade. We publish all of it because none of it is the difference.

The difference is two modules called Discovery and Evolution. Discovery decides what a question actually is and who should be in the room to answer it. Evolution decides what the room got wrong and what changes next time. Between them they are the reason the same models, available to anyone at the prices listed above, produce a different result for us than they will for you.

That is the whole description you are going to get. Not the architecture, not the seat bindings, not the taxonomy underneath either, and not the loop that connects them. We would rather be measured on the numbers we publish than admired for a diagram — and every figure on this page stands on its own without knowing a thing about either module.

The conflict, disclosed

Where is Anthropic in all this?

Eighteen models, four clouds, and almost no Claude. That is not an oversight and it is not a verdict on the models.

Claude is not on the bench because Claude is the operator. It is the primary controlling interface for this system — it reads the state, writes the specs, dispatches the work, and reviews what comes back. Something that decides what gets measured cannot also compete in the measurement. We removed Anthropic's models from panel duty in July for exactly that reason, and the scoreboard has stayed clear of the scorer since.

Two narrow exceptions survive. Claude Haiku is the dedicated bench sub: when a panel seat exhausts every retry and every region, Haiku comes off the bench and takes the seat once, so the panel is never silently short a voice. It is stamped as a substitution every time it plays and counts toward no accuracy figure here — a substitute's minutes do not go on the starter's average. And Claude Fable is used sparingly by rule, reserved by a standing cost rule for one narrow tier of the hardest work and barred from the high-traffic paths.

If that arrangement gives us a bias, it points the opposite way from the obvious guess: the company whose model runs our control plane is the one family we refuse to let compete.

Method

How this was measured, and what it does not show

What we did

Items. 100 public test items — 50 GSM8K test, 50 Banking77 test — run twice on two seeds. Every candidate saw the identical prompt under the same output rules.

Calls. 1,800 live calls per run, 3,600 in total. No cached answers, no fixtures. Panel comparisons are simulated from the recorded grid, so no model got a second chance the others did not.

Route and identity. Every row names the cloud the model was called on, and every call records the model the provider reports serving. We refuse to score a seat whose served model is unknown.

What it does not show

Prices are advisory. List-price estimates for a roughly 1,000-token exchange at each cloud's published rate. Estimates, not invoices, and list price rather than what any particular buyer pays.

The limits. Two runs, one day, public test splits, one prompt schema. This measures short-answer accuracy under our prompt rules. It is not a general capability ranking and we would not defend it as one.

Replication moved the numbers. An earlier 40-item run flattered the leaders by about four points, and one result did not survive a second seed at all. Both are printed rather than dropped.

If you think a number here is wrong, the useful thing to send us is which item and which model — we will rerun it and print what happens. That is how the Llama section above came to exist.