Case study · the comparison everyone is making

We asked 22 AI models to compare the AI boom to the dot-com bubble. They said the comparison is wrong.

The mandate asked for ten named scenarios, each priced, across four horizons. The panel delivered them — and ranked Dot-Com Replay dead last at 1%. Its own evidence layer then flagged the premise it had been handed as CONTRADICTED. One model held a REJECT at 95% confidence through all three debates, arguing the whole exercise was invalid. We printed that too.

229model calls
22AI models · 3 clouds
3debate panels · 43 seats
24m 44stotal run time
29/30requested forecasts answered
Want to hear Rex, our AI influencer, explain it?
Ten scenarios, priced. Dot-Com Replay finished last.
The contract that forced every scenario to be answered, the 65% near-term correction odds, and the one model that never agreed.
Act 1 · the setup

A thirty-item contract, not a question

Most "is AI a bubble" analysis is an essay. This was a contract: ten scenarios by name, twenty numbered questions, and four horizons — 12 months, 24 months, 3 years, 5 years — with every forecast required to carry its assumptions, its strongest counter-argument, the indicators that would move it, and a reassessment date.

That mattered because of what happened the first time. An earlier run of the same mandate answered 14 of 30 items. Every named scenario was missing. Nothing had been truncated — the brief was built at 30,950 characters. The system had read the whole thing and quietly compressed it, because a hardcoded instruction capped scenario output at "2–4 outcomes" regardless of what was asked for.

We built a coverage contract to stop that: the mandate is parsed into an explicit register of every requested item, the register is injected into the brief, and coverage is measured against it afterwards. This run answered 29 of 30. The one miss is named below rather than hidden.

Why this is published as method, not prophecy. These are the panel's estimates, dated so they can be scored against reality later. They are model-generated analysis — not investment advice, and not a recommendation to buy or sell any security.
Act 2 · the answer

Ten scenarios, priced

Each scenario was named in the mandate. The panel assigned probabilities summing to 100 and was not permitted to merge or drop any of them.

ScenarioProbabilityRank
Persistent Expansion30%1
Orderly Consolidation25%2
Hyperscaler Oligopoly20%3
Open-Weight Commoditization10%4
Infrastructure Reckoning5%5
Government-Controlled AI3%6=
Fragmented Multipolar AI3%6=
Small-Model Disruption2%8
Agent Economy Breakthrough1%9=
Dot-Com Replay1%10

The scenario the entire mandate was built around finished last. Asked to price a repeat of 1999 — valuations collapse, most startups fail, infrastructure survives, a few winners emerge — the panel put it at 1%. The top three outcomes, totalling 75%, all describe a market that consolidates rather than collapses.

The one-sentence call

A 20–30% pullback in AI startup valuations will happen within 18 months. The delivered report, page one — 76 characters, no hedging

Underneath it, the near-term picture is sharper than the scenario table suggests: the research layer put a 65% probability of a ≥30% correction in AI-related equities within 12 months, and 55% on a ≥50% decline within 24 months. A correction is likely. A dot-com replay is not. Those are different claims, and the distinction is the finding.

Open versus closed, by 2029

Open-weight models take 40–45% of usage. Closed models retain 55–60% of economic value. The panel forecast share rather than declaring a winner — the mandate explicitly forbade normative "which is better" claims.

Act 3 · the dissent

One model never agreed. It is in the report.

Across three panels the vote moved 9–3, then 11–2, then 15–1. Dissent narrowed but never disappeared. Qwen3 held REJECT at 95–98% confidence from the first debate to the last, and its objection was not about the numbers — it was about whether the question should be answered at all.

The dot-com bubble and the AI investment cycle are ontologically distinct systems. Publishing calibrated probabilities under this analogy does not illuminate uncertainty; it creates the illusion of precision where none can exist. Falsifiability without structural validity is not rigor — it is ritualized misdirection. Qwen3 — REJECT at 95% confidence, sustained across all three panels
Publishing under continuity assumptions risks giving investors and policymakers an incomplete risk surface. Grok 4.3 — REJECT at 68% confidence, panel two

Five analysts changed position during the debates. Qwen3 was not one of them. A system that only reports its majority would have shown you a 15–1 consensus and never mentioned that its most confident participant thought the exercise was a category error.

Act 4 · the evidence test

It labelled its own premise CONTRADICTED

Every material claim carries an evidence label. The one that matters here is the mandate's own founding assumption:

ClaimLabel
"AI investment cycle shares similarities with dot-com bubble" — the premise the report was built onCONTRADICTED
Amazon, Microsoft, Alphabet, Apple survive consolidationINFERRED
AI technologies retain long-term economic value despite correctionsINFERRED
65% probability of ≥30% AI equity correction in 12 monthsASSUMED
Open-weight models capture 40–45% of usage by 2029ASSUMED

Zero claims were labelled VERIFIED, and we are not hiding that. This is a forecast about a future that has not happened; almost nothing in it can be verified today. A report of this kind that claimed a wall of verified facts would be lying about its own epistemics. What it can do is mark clearly which parts rest on evidence, which on inference, and which on assumption — and it does.

Confidence: MODERATE. 75% panel agreement, 16 analysts on the final panel, 5 documented position changes, confidence spread 0.2. The system did not round its own uncertainty away.

Questions this case answers directly

Plain answers to what people actually ask. Every figure comes from the delivered report for case 2026-9410 — a 22-model adversarial panel across 3 debate rounds. These are clearly-labelled panel estimates, dated for later scoring — not investment advice, and not a recommendation regarding any security.

Is the AI market about to crash like the dot-com bubble?

Probably not like the dot-com bubble — but a correction is more likely than not. A 22-model adversarial panel run by 3Dogs Nexus put Dot-Com Replay at 1%, the lowest of ten named scenarios, while simultaneously putting 65% on a ≥30% correction in AI-related equities within 12 months and 55% on a ≥50% decline within 24 months. Those two findings only look contradictory if you treat "correction" and "dot-com replay" as the same event. The panel's position is that the AI cycle corrects without collapsing: enterprise adoption is stickier than the consumer-driven boom of 1999, and the infrastructure layer is controlled by companies with real cash flows.

How does the AI bubble compare to the dot-com bubble?

Less closely than the comparison's popularity suggests — and the panel said so about its own mandate. Asked explicitly to compare the two, a 22-model panel labelled the claim "AI investment cycle shares similarities with dot-com bubble" as CONTRADICTED by the evidence available to it. The structural differences it identified: AI's deflationary compute trajectory, enterprise-embedded rather than consumer-speculative adoption, vertical integration by incumbents with existing revenue, and open-weight commoditization with no 1999 analogue. The dissenting seat went further, calling the two "ontologically distinct systems."

When will the AI bubble burst?

The panel's dated answer: a 20–30% pullback in AI startup valuations within 18 months, with a 65% probability of a ≥30% equity correction inside 12 months. It declined to forecast a "burst" in the 1999 sense at any horizon, pricing that scenario at 1%. The distinction it insists on is between a re-rating — valuations falling to meet revenue — and a collapse in which the underlying technology loses economic relevance. It considers the first likely and the second remote.

What percentage of AI startups survive?

Survival concentrates by category rather than spreading evenly. The panel found that firms demonstrating profitability, diversified revenue, real-world problem-solving and operational adaptability are the ones that persist, and that infrastructure leaders historically capture the post-correction market — benchmarks from previous consolidations suggest survivors control 70–80% of the resulting market. Note the honest gap: this run did not deliver a clean split between outright failure and acquisition, which was one of the thirty requested items and is disclosed as the single miss below.

Will open-weight or closed models win?

Neither, and the framing is the problem. By 2029 the panel forecasts open-weight models taking 40–45% of usage while closed models retain 55–60% of economic value — volume and value separating rather than one side winning. The mandate explicitly barred normative "which is better" claims, so the panel forecast share instead. Any source telling you one model type has simply won is answering a different, easier question.

What is the risk of stranded data-centre assets in AI?

Lower than the alarm suggests, on this panel's reading: Infrastructure Reckoning — the scenario in which data-centre and energy spending overshoots real demand, forcing write-downs and distressed sales — was priced at 5%, eighth of ten. The reasoning is that AI infrastructure is being bought by a small number of counterparties with substantial existing cash flows, unlike the debt-financed telecom overbuild of 1999–2001, and that compute demand is deflationary in price but not in volume.

What we got wrong, and what we are not claiming

One requested item went unanswered. Of thirty, the panel did not deliver a clean percentage split between AI companies that fail outright and those absorbed by acquisition or merger. It is disclosed here rather than quietly dropped — that disclosure is the entire point of measuring coverage.

One horizon came back thin. The 3-year band is less well populated than 12 months, 24 months and 5 years.

Nothing is VERIFIED. Zero of the material claims carry a VERIFIED label, because this forecasts a future that has not occurred. Every probability is an estimate with a date attached, published so it can be scored later and marked wrong if it is wrong.

And the strongest argument against publishing this at all is in the report. Qwen3's position — that calibrated probabilities on a structurally mismatched analogy manufacture false precision — was never refuted, only outvoted. We think it is worth reading. That is why it is printed rather than averaged away.

The delivered report

Case 2026-9410 in full: the one-sentence call on page one, the ten scenario probabilities, the evidence classifications, the preserved dissent, and the complete panel record. 229 model calls · 22 AI models · 24m 44s.

Open the full report (PDF)

Run this on your own question

This case was a self-run demonstration on the 3Dogs Nexus development system, not a paid engagement. The method is the product: independent models from different vendors argue your decision, disagreement is preserved rather than resolved, and every claim comes back labelled verified, inferred, assumed or contradicted.

Start a decision case More case studies
Watch on YouTube →