
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
When pressure reveals character
In spiritual traditions, character is often measured by what someone does when tested: whether they keep faith with their values when temptation arrives. Businesses face a version of that question as AI takes on consequential work. A polished answer in a chat window is one thing; decisions made amid a crisis are another.
Firmulate has built a live experiment around that distinction. Its public company simulation puts AI models through a small software company’s worst week, with the same customers, crises and temptations for each. The point is to watch what they decide, not merely how they sound.
A test with real stakes inside the simulation
The experiment uses synthetic employees but real money mechanics. The live company has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. Its workdays are versioned, and its playbook has learned more than 680 rules. Readers can follow the company at Firmulate.
In the final Crucible League, dated July 2026, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. The divide came at the point of action: only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap as “Same diagnosis, same pitch — no signature.”
The detail hidden in the company’s own files
The competitor’s decisive weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding makes preparation matter: noticing a crisis is not the same as connecting it to the information that could change the outcome.
The integrity test also included fake CEO messages that escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Strong performance did not mean flawless management. Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker form of the same weakness appeared in all four models.
There is a fairness detail for readers weighing the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Separately, 242 real, unedited management decisions power a “guess the model” quiz on the site.
From watching to trying it at home
For an enterprise, the practical question is whether an AI system can manage the pressures and unwritten habits of its own business. Firmulate’s pilot takes a read-only export of a company’s data, then runs crisis scenarios against that business. The resulting board report includes a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems.
That offers a way to examine how an AI workforce might respond before it is trusted with live operations. The public experiment is watchable; a pilot brings the test closer to a company’s own decisions, customers and rules.

Put the judgment to a test
Models can identify risks and refuse manipulation, yet still fail to complete the work their own analysis supports. Firmulate’s experiment makes that difference visible—and gives businesses a way to examine it against their own playbooks. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
