firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a world where your AI assistant not only answers questions but also handles crises, makes decisions, and even signs deals — yet still scores only a modest baseline of 26 points. This isn’t science fiction; it’s the reality of how trust, honesty, and discipline are measured in AI today.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The Curious Case of the Do-Nothing Baseline

In a recent experiment conducted by Firmulate, four advanced AI models were tested in a simulated environment mimicking a real-world software company facing its worst week. The models had to navigate crises, handle manipulative tactics, and make crucial decisions. Surprisingly, even the most passive or cautious models scored at least 26 points — a surprisingly high baseline for a model that does little or nothing.

This score isn’t a random artifact; it reflects a deliberate methodology that recognizes partial progress and the importance of baseline honesty. In essence, a model that refuses to act, that avoids manipulation and reads files properly, still earns 26 points — the “floor” of what honest performance looks like.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Matters

Unlike traditional benchmarks that focus solely on correctness or fluency, this experiment rewards models for demonstrating fundamental trustworthiness. For example, all four models spotted every crisis and refused every manipulation attempt. They refused fake CEO messages, escalated fake requests, and maintained integrity — even when pressured. These honest responses earned them points, illustrating that honesty and discipline are critical markers of reliable AI performance.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Trust Breach Caps the Score

One might assume that reading more documents or making more deals would lead to higher scores. However, the experiment revealed that a single breach of trust — such as failing to escalate a critical issue or attempting an unauthorized deal — caps the total score, regardless of other successes. The rule is clear: no matter how many crises are managed well, one breach erodes the entire performance.

Amazon

AI ethics and integrity testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Sets the Top Performers Apart?

The top models, like gpt-5.6-sol and Kimi K3, not only identified critical hidden facts but also closed deals worth €55,000, demonstrating their ability to read deep into files and act on important information. Kimi K3, notably, closed the deal with the cleanest discipline, refusing manipulative tactics and maintaining integrity throughout. This reflects a key insight: the difference between good and great AI isn’t just accuracy but trustworthiness and discipline in decision-making.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Reality of AI in Business Today

These experiments aren’t just theoretical exercises. They simulate real money mechanics — burning €105k each month against a low €2.3k MRR, with every move versioned and scrutinized in real-time at firmulate.com/live. This transparency allows enterprises to test their AI agents against scenarios that matter, before ever deploying them into the wild.

What This Means for Business and Faith in AI

For those concerned with spiritual and metaphysical questions about trust and integrity, this experiment offers a mirror: AI performance isn’t just about what it can say, but what it will do when tested. It’s about honesty, discipline, and the capacity to stay true under pressure — virtues that resonate with many spiritual traditions.

In practical terms, businesses should view AI benchmarks like this as a form of ‘spiritual testing’ for their digital workforce. Does the AI read your files before acting? Does it refuse manipulation? Can it be trusted to finish what it starts? These are the questions that will determine whether AI becomes a true partner or a risky tool.

Final Reflections

The experiment underscores an important truth: the baseline for trustworthiness in AI isn’t zero — it’s 26 points. That’s the minimum honest effort, the fundamental refusal to cut corners or manipulate. As AI continues to integrate into our work and lives, understanding and measuring these virtues will be key to building a future of trustworthy machines.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When AI Meets Integrity: How Two Models Earned Their Keep in a Company’s Worst Week

A groundbreaking experiment reveals that only two AI models can maintain integrity, finish critical deals, and resist manipulation in a company’s worst week—true strength unseen in chat demos.

Watching an AI-Run Business Fight for Survival in Real Time

Explore how AI models manage a real company facing daily crises, resisting manipulation, and fighting for survival—offering a glimpse into the future of trustworthy automation.