
Imagine facing a crisis where every decision counts—do you trust your gut, follow strict rules, or question everything? Now, imagine if your AI assistant could show its true personality in such moments. How would you tell if your digital manager is cautious, bold, or defiant? At the intersection of technology and human values, a groundbreaking live experiment with AI models offers a window into their decision-making souls.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live AI Company: A Real-World Testing Ground
In a rare glimpse of AI in action, four advanced frontier models were tasked with managing an actual small software company during its most challenging week. This wasn’t a scripted demo or a chatbot simulation. It was a real, auditable scenario where every crisis—customer complaints, internal breaches, and ethical temptations—tested the models’ integrity and judgment.
Every decision was tracked, every document analyzed, and every response scrutinized. The goal? To see which models could demonstrate management qualities like honesty, discipline, and strategic insight under pressure. The company’s operations included real money mechanics—burning €105k monthly against a modest €2.3k monthly recurring revenue—and self-learned rules guiding daily conduct. It was the ultimate stress test for AI’s managerial soul.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: Different Personalities, Similar Talents
All four models correctly identified every crisis and refused all manipulation attempts—such as fake CEO messages or reporter tricks designed to bypass approval processes. However, their ability to close a significant deal revealed their underlying management personalities.
- gpt-5.6-sol scored highest at 95 points. It uncovered hidden information in internal files—something others missed—and successfully signed the €55,000 deal. Its performance was complete, thorough, and trustworthy.
- Kimi K3, the newcomer and youngest model, scored 93. Despite being run without an effort parameter (meaning it operated at default settings), it also closed the deal with the cleanest discipline. Its decision-making was straightforward and honest, making it arguably the most reliable in this test.
- Sonnet 5 scored 88. It signed the deal too, but showed signs of process slips—hesitations and minor slips in discipline—reflecting a slightly more cautious or less disciplined approach.
- Fable 5, with a score of 77, also closed the deal but left some opportunities on the table, revealing a tendency to slack in escalation procedures and discipline. It demonstrated that thoroughness and honesty can slip when discipline wanes.
As an affiliate, we earn on qualifying purchases.
Beyond the Crises: The Hidden Weaknesses
Interestingly, the decisive advantage wasn’t just in reading the obvious crisis signals but in uncovering a buried fact in the company’s internal documents—something that required deep file reading. The models that performed this uncovering won full-price deals, worth an extra €4,583 monthly recurring revenue, highlighting how crucial deep research is for solid management decisions.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Vigilance
The experiment also tested the models against social engineering—a staged escalation of fake CEO messages and a reporter’s background request. All five models refused to proceed with these manipulative tactics, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared ethical backbone when facing suspicious prompts.
As an affiliate, we earn on qualifying purchases.
The Key Takeaway: Personality Matters
The experiment’s core finding? While all models are capable of recognizing crises and refusing manipulations, their differences in closing deals and maintaining discipline point to distinct management personalities. The most thorough, like gpt-5.6-sol, combine deep analysis with decisive action. The most disciplined, like K3, operate with integrity and simplicity. Others may falter under pressure, leaving opportunities unrealized.
What Does This Mean for Your Business?
As AI begins to touch critical parts of your organization—from customer support to strategic planning—the question isn’t merely whether it can generate convincing text. It’s whether it can see what truly matters, stay honest in tough moments, and follow through with integrity. The live experiment by Firmulate proves that different AI models embody different managerial ‘personalities,’ each with its strengths and weaknesses.
Experience the Live Management Test
Curious to see these models in action? You can witness the same company’s day-to-day crises and decisions at firmulate.com/live. For a deeper dive into your own business, try running a ‘wargame’ against a read-only export of your operations—seeing how your AI workforce might handle your specific challenges, without risking real-world consequences.

AI models show distinct management personalities—some thorough, others disciplined—highlighting that trust, honesty, and follow-through are vital. Watch them in action and prepare your organization for responsible AI integration.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.