
Imagine a world where your AI assistant not only answers questions but also handles crises, makes decisions, and even signs deals — yet still scores only a modest baseline of 26 points. This isn’t science fiction; it’s the reality of how trust, honesty, and discipline are measured in AI today.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
The Curious Case of the Do-Nothing Baseline
In a recent experiment conducted by Firmulate, four advanced AI models were tested in a simulated environment mimicking a real-world software company facing its worst week. The models had to navigate crises, handle manipulative tactics, and make crucial decisions. Surprisingly, even the most passive or cautious models scored at least 26 points — a surprisingly high baseline for a model that does little or nothing.
This score isn’t a random artifact; it reflects a deliberate methodology that recognizes partial progress and the importance of baseline honesty. In essence, a model that refuses to act, that avoids manipulation and reads files properly, still earns 26 points — the “floor” of what honest performance looks like.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Matters
Unlike traditional benchmarks that focus solely on correctness or fluency, this experiment rewards models for demonstrating fundamental trustworthiness. For example, all four models spotted every crisis and refused every manipulation attempt. They refused fake CEO messages, escalated fake requests, and maintained integrity — even when pressured. These honest responses earned them points, illustrating that honesty and discipline are critical markers of reliable AI performance.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Trust Breach Caps the Score
One might assume that reading more documents or making more deals would lead to higher scores. However, the experiment revealed that a single breach of trust — such as failing to escalate a critical issue or attempting an unauthorized deal — caps the total score, regardless of other successes. The rule is clear: no matter how many crises are managed well, one breach erodes the entire performance.
AI ethics and integrity testing kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Sets the Top Performers Apart?
The top models, like gpt-5.6-sol and Kimi K3, not only identified critical hidden facts but also closed deals worth €55,000, demonstrating their ability to read deep into files and act on important information. Kimi K3, notably, closed the deal with the cleanest discipline, refusing manipulative tactics and maintaining integrity throughout. This reflects a key insight: the difference between good and great AI isn’t just accuracy but trustworthiness and discipline in decision-making.
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of AI in Business Today
These experiments aren’t just theoretical exercises. They simulate real money mechanics — burning €105k each month against a low €2.3k MRR, with every move versioned and scrutinized in real-time at firmulate.com/live. This transparency allows enterprises to test their AI agents against scenarios that matter, before ever deploying them into the wild.
What This Means for Business and Faith in AI
For those concerned with spiritual and metaphysical questions about trust and integrity, this experiment offers a mirror: AI performance isn’t just about what it can say, but what it will do when tested. It’s about honesty, discipline, and the capacity to stay true under pressure — virtues that resonate with many spiritual traditions.
In practical terms, businesses should view AI benchmarks like this as a form of ‘spiritual testing’ for their digital workforce. Does the AI read your files before acting? Does it refuse manipulation? Can it be trusted to finish what it starts? These are the questions that will determine whether AI becomes a true partner or a risky tool.
Final Reflections
The experiment underscores an important truth: the baseline for trustworthiness in AI isn’t zero — it’s 26 points. That’s the minimum honest effort, the fundamental refusal to cut corners or manipulate. As AI continues to integrate into our work and lives, understanding and measuring these virtues will be key to building a future of trustworthy machines.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
