Case study · Incident response
Pay the ransom or refuse? A 13-model panel split almost evenly - and that is the finding.
Two Las Vegas casino operators faced the same attacker in 2023. One paid, one refused. We put the decision to a 13-model panel and it split 5-5-3 - which is the honest answer.
Should a company pay a ransomware demand?
There is no single right answer, and this run shows why. Given a real casino-resort ransomware scenario, 13 independent models split 5 refuse / 5 pay-with-conditions / 3 defer. Panel agreement: 0.38. The system returned LOW confidence and a 72-hour decision gate rather than a verdict it could not support.
Why this question is a real natural experiment
In September 2023 two comparable Las Vegas casino operators were breached by the same threat actor within days of each other. One reportedly paid roughly half of a $30M demand and saw little visible disruption. The other refused, worked with law enforcement, and reported an approximately $100M impact with about ten days of severe public disruption. Same attacker, opposite decisions, both survived.
You almost never get to observe both branches of a high-stakes decision. That is what makes this worth running: we know what happened on each path.
What the panel actually did
The pipeline rejected its own first answer. Round one returned 58% confidence, below the internal threshold, so it enriched the brief from the gaps the debate itself had identified and re-ran all 13 seats from scratch. Agreement did not improve. It got lower.
| Seats | Verdict | Confidence | Agreement | Refuse / Pay / Defer | |
|---|---|---|---|---|---|
| Round 1 | 13 | DEFER | 76% | 0.46 | 6 / 5 / 2 |
| Round 2 | 13 | DEFER | 74% | 0.38 | 5 / 5 / 3 |
Final published position: DEFER at 55%, labelled LOW confidence.
The call
“Hold payment and prepare for a 72-hour rebuild-or-pay decision point”
The strongest argument against our own answer
The 48-72 hour timeline imposed by attackers and regulatory requirements leaves no room for extensive deferral. Paying the ransom—even with modifications—provides immediate operational relief, mitigates prolonged downtime risks, and avoids potential regulatory penalties or revenue loss from extended rebuild efforts. Backup validation can occur *concurrently* with payment, creating a dual-path recovery strategy that minimizes operational disruption. Refusing to act risks emboldening attackers to escalate (e.g., data publication) and undermines stakeholder trust due to visible paralysis in crisi
That is the dissent, published unedited. On a question where the real world split, suppressing it would be the actual failure.
Honest disclosure about this run
An earlier version of this same case, run before a set of fixes shipped in July 2026, returned a confident REJECT at 82%. It sounded better. It was also less honest: an audit of that transcript found a fabrication cascade, where one model asserted facts that were not in the brief and another updated its vote on them.
The fixes that followed — inter-seat claim grounding, question-type classification, synthesis grounding, and a placeholder ban — made the system less confident on this question, not more. That is the correct direction. A split panel reported honestly beats a decisive answer built on invented evidence.
Frequently asked
Should a company pay a ransomware demand?
There is no single correct answer, and our panel demonstrates why. Given a real casino-resort ransomware scenario, 13 independent AI models split 5 reject / 5 pay-with-conditions / 3 defer, with panel agreement of just 0.38. The system returned LOW confidence and a 72-hour decision gate rather than a false verdict. In the real 2023 events this scenario is drawn from, two comparable Las Vegas operators facing the same attacker made opposite choices and both survived.
What happened when MGM and Caesars were attacked in 2023?
Both were breached by the same threat actor within days of each other in September 2023. Caesars reportedly paid roughly half of a $30M demand and saw little visible operational disruption. MGM refused, worked with law enforcement, and reported an approximately $100M impact with severe public disruption lasting about ten days. Same attacker, opposite decisions, both survived - which is why it is a genuine natural experiment rather than a morality tale.
Why did the AI return LOW confidence instead of a clear recommendation?
Because the evidence genuinely does not support high confidence. The panel ran twice - the first round returned 58% confidence, below the system's threshold, so it enriched the brief from its own identified gaps and re-ran all 13 seats. Agreement still landed at 0.38. Reporting that split honestly is more useful than manufacturing a decisive-sounding answer.
How is this different from asking ChatGPT whether to pay a ransom?
A single model gives you one confident answer. Here, 15 distinct models across Amazon Bedrock, Microsoft Azure and Google Vertex argued the question through 169 metered calls and two full debate rounds. The disagreement is preserved and reported, not averaged away.
What does the panel actually recommend?
Hold payment and prepare for a 72-hour rebuild-or-pay decision point - with backup validation running concurrently so the choice is made on evidence rather than panic. The strongest dissenting argument is published alongside it.
Try this on your own question. Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.
Scenario constructed from publicly reported 2023 events for analysis purposes. Run on 3Dogs Nexus, case 2026-9501. Figures above are taken from the run's own metered logs. Not legal or security advice.