# Six models, fifty problems, four clouds

Measured 5 September 2026 · 300 live API calls per run · 50 original problems · 6 models · 4 clouds · no cached or fixture responses. Run receipts: https://3dogs.ai/model-reviews/six-models-fifty-problems/receipts.json

The most expensive model in the panel costs 3.9× as much per correct answer — cost divided by accuracy, not the per-token price — as one that scores within two points of it. The cheapest is 8× cheaper again and gets a third of the questions right. This is the teardown 3Dogs runs on a client's own prompts, shown on fifty problems written for this page.

## Results (Run A)

$ per million correct answers = cost per million answers ÷ accuracy. Costs are estimates from published per-token prices, tokens estimated at four characters per token. Published prices move by region and promotion. Estimates, not invoices.

| Model | Cloud | Accuracy | Correct | $/1M answers | $/1M correct | Median latency |
|---|---|---|---|---|---|---|
| GPT-5.5 | Azure | 88% | 44/50 | 45.63 | 51.85 | 1.5s |
| Gemini 3.8 Flash | Google Vertex | 86% | 43/50 | 11.45 | 13.32 | 3.7s |
| Qwen3 Next 80B | Google Vertex | 56% | 28/50 | 4.74 | 6.77 | 0.5s |
| DeepSeek V3.2 | Azure | 50% | 25/50 | 7.27 | 14.54 | 0.6s |
| Llama 3.3 70B | AWS Bedrock | 46% | 23/50 | 20.09 | 43.67 | 0.3s |
| Nova Lite | AWS Bedrock | 32% | 16/50 | 1.94 | 6.06 | 0.5s |

Qwen3 Next returned nothing on 10 of its 50 calls; accuracy is scored over all 50.

## The cheap-model advantage reverses when the task gets harder

On an easier twenty-item set run the day before, Nova Lite matched GPT-5.5 on accuracy at a fraction of the price. On these fifty multi-step problems it gets 32% and GPT-5.5 gets 88%. Same models, same week, opposite conclusion. A benchmark you did not write cannot tell you which model to buy; only your own prompts can.

## Our own instrument was wrong twice before these numbers were right

The hand-written answer key had two errors and one ambiguous question, all found by 3Dogs' own answer-key audit before any model was scored (two independent programs per item; a key is flagged only where both ran, agreed with each other, and disagreed with the key). Separately, an earlier scorer marked a model wrong for answering 7.17 when the exact value was 7.1666…; a number written to two decimals claims accuracy to half a unit in its last place, and the scorer now honours that.

## Reproducibility: Run B beside Run A

| Model | Run A | Run B | Change |
|---|---|---|---|
| GPT-5.5 | 88% | 88% | — |
| Gemini 3.8 Flash | 86% | 86% | — |
| Qwen3 Next 80B | 56% | 56% | — |
| DeepSeek V3.2 | 50% | 50% | — |
| Llama 3.3 70B | 46% | 52% | +6 pts |
| Nova Lite | 32% | 36% | +4 pts |

## Method, and what this does not show

50 original multi-step word problems with a single numeric answer, published in the receipts so anyone can re-run them; future issues use fresh problems. Six models on four clouds, 300 live calls per run, two runs, identical prompt, no cached answers; every call records the model the provider reports serving. A numeric answer is correct if it matches the key to the precision it was written to; an empty or unparseable answer is wrong. This measures multi-step arithmetic under one prompt rule set on one day; it is not a general capability ranking. Your prompts change the table — that table is the product. Measure on your own prompts ($499): https://3dogs.ai/pricing/

Machine-readable: https://3dogs.ai/model-reviews/six-models-fifty-problems/receipts.json
