firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

When pressure reveals character

In spiritual traditions, character is often measured by what someone does when tested: whether they keep faith with their values when temptation arrives. Businesses face a version of that question as AI takes on consequential work. A polished answer in a chat window is one thing; decisions made amid a crisis are another.

Firmulate has built a live experiment around that distinction. Its public company simulation puts AI models through a small software company’s worst week, with the same customers, crises and temptations for each. The point is to watch what they decide, not merely how they sound.

A test with real stakes inside the simulation

The experiment uses synthetic employees but real money mechanics. The live company has 13 employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. Its workdays are versioned, and its playbook has learned more than 680 rules. Readers can follow the company at Firmulate.

In the final Crucible League, dated July 2026, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. The divide came at the point of action: only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap as “Same diagnosis, same pitch — no signature.”

The detail hidden in the company’s own files

The competitor’s decisive weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding makes preparation matter: noticing a crisis is not the same as connecting it to the information that could change the outcome.

The integrity test also included fake CEO messages that escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong performance did not mean flawless management. Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker form of the same weakness appeared in all four models.

There is a fairness detail for readers weighing the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Separately, 242 real, unedited management decisions power a “guess the model” quiz on the site.

From watching to trying it at home

For an enterprise, the practical question is whether an AI system can manage the pressures and unwritten habits of its own business. Firmulate’s pilot takes a read-only export of a company’s data, then runs crisis scenarios against that business. The resulting board report includes a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems.

That offers a way to examine how an AI workforce might respond before it is trusted with live operations. The public experiment is watchable; a pilot brings the test closer to a company’s own decisions, customers and rules.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the judgment to a test

Models can identify risks and refuse manipulation, yet still fail to complete the work their own analysis supports. Firmulate’s experiment makes that difference visible—and gives businesses a way to examine it against their own playbooks. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI’s Hidden Files Decide Your Business Fate: The Power of Reading Deeper

Deep reading AI models can uncover hidden business facts buried two files deep, influencing deals and trust. The future of AI relies on understanding beneath the surface.

The ‘More Passes’ Mistake in At-Home Facial Techniques

With more passes or layers, your skin may suffer; discover how to prevent this common at-home facial mistake and protect your skin’s health.

The Difference Between Skin Stimulation and Skin Stress

Understanding the difference between skin stimulation and stress can help you protect and enhance your skin’s health—discover how to do it effectively.

Why Conductive Layers Matter More Than Most People Realize

Ineffective conductive layers can cause device failures, and understanding their importance reveals how crucial they truly are for reliable electronic performance.