
Imagine testing a new gardening tool that should help you grow better plants, but instead, it always scores at least 26 out of 100, even when you do nothing. Sounds odd, right? That’s exactly what AI benchmarking reveals about how we judge the honesty and reliability of artificial intelligence systems.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Do-Nothing Baseline
When evaluating AI models, especially for business use, it’s essential to know how they perform in the worst-case scenario — doing nothing. Surprisingly, even a totally inactive AI model scores 26 points in a recent live benchmark. Why isn’t it zero? Because even the simplest models can produce partial, trivial outputs that count as a form of progress.
Why Partial Progress Counts
In this benchmark, every small decision or piece of information that the AI correctly handles adds to its score. For example, if an AI recognizes a crisis but doesn’t act on it, it still gains some points — reflecting that it’s at least aware of the situation. This approach ensures the scoring system rewards genuine understanding, not just superficial chatter.
The Impact of Trust Breaches
However, the scoring system also has built-in safeguards. If an AI model breaches trust — for instance, by attempting manipulation or ignoring critical data — its total score is capped. This means no matter how well it performs elsewhere, a single breach prevents it from earning a perfect score, mirroring real-world importance of integrity in AI systems.
AI decision-making assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Testing AI in a Simulated Business Crisis
Firmulate ran four leading AI models through a detailed simulation of a small software company facing its worst week. The scenario included actual customer crises, internal errors, and ethical dilemmas. Each model was tasked with making management decisions in a fully auditable environment, mimicking real business pressures.
Results Highlighted Honesty and Competence
All models correctly identified every crisis and refused manipulation attempts, including fake CEO messages and subtle bribery tactics. Notably, only two models signed a €55,000 deal based on their analysis — the true measure of their effectiveness. The other two recognized the opportunity but did not act decisively, leaving money on the table.
The Hidden Weaknesses Discovered
Interestingly, the decisive advantage often lay in reading deeper into company documents. The models that examined files thoroughly uncovered critical information that wasn’t immediately visible elsewhere — and as a result, they closed the deal at full price, adding an estimated +€4,583 MRR to the company’s revenue.
enterprise AI ethics and trust software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust and Discipline Matter in AI
During social engineering tests, all models refused requests to bypass approval or impersonate executives, demonstrating strong ethical boundaries. Yet, some struggled with discipline — for example, Opus 4.8’s most thorough participant failed to escalate issues properly, leaving opportunities on the table despite its detailed analysis. That highlights that even the best models can falter without consistent discipline.
AI benchmarking and evaluation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business Leaders
For companies considering AI for management or decision-making, the key takeaway isn’t just how well an AI writes or talks — it’s whether it can finish what it starts, interpret data responsibly, and remain honest under pressure. A model’s real value lies in its ability to handle complex, ethically charged scenarios without shortcuts.
The Benchmark’s Transparency and Fairness
Firmulate’s live experiments are openly accessible, with decision-making processes fully versioned and auditable. This approach ensures that evaluation is both fair and transparent, avoiding the common pitfall of inflated scores based on superficial capabilities.
AI model testing simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Thoughts: An Honest Measure of AI Readiness
While AI models continue to improve, these benchmarks reveal that even in their best form, they must be held accountable for honesty and thoroughness. The starting score of 26 points for a do-nothing model underscores that partial progress and ethical integrity are critical to truly valuable AI systems. As AI integrates deeper into business operations, understanding these nuances becomes vital for success.
Explore the Benchmarks
Curious about how various models compare? Visit Firmulate’s benchmarks page for full results and insights into this ongoing experiment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
