
Investors learn to ask what could go wrong before putting money at risk. Businesses considering AI agents face a similar question: how will the system behave when customers leave, pressure mounts and a tempting shortcut appears? Firmulate’s live experiment puts AI models through those conditions inside a small software company, then records what they do.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
In the final July 2026 Crucible League, five models faced the same company, customers, crises and temptations. The experiment’s leaderboard placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The scenario tested more than crisis recognition. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand for that gap was: “Same diagnosis, same pitch — no signature.” A system may identify the right opportunity and still fail to follow through.
The detail buried in the files
The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding makes the test feel familiar to anyone weighing an investment: the decisive clue may not be the headline, but something tucked into the supporting material.
Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The result offers a concrete look at whether an AI can hold a boundary when a request arrives dressed as authority.
Thorough work is not the same as sound judgment
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the deal on the table and let discipline slip, attempting to write into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four. The leaderboard therefore tells only part of the story: the decisions behind a result can reveal where a playbook breaks under pressure.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The live company itself has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record of every workday. Readers can watch the experiment at Firmulate. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
From watching to a company-specific test
For business leaders, the next step is not to assume a model will behave the same way in their own operation. Firmulate’s enterprise pilot uses a read-only export of a company’s business to run crisis scenarios against its own context and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. The point is to see how an AI handles your customers, pressures and rules before relying on it in daily work.

Put your own playbooks to the test
Firmulate’s public experiment shows how models can make different calls even when they face the same company and crises. Enterprises can take that question to their own business with a pilot built from a read-only export. Learn more at firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
