firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI for your busy restaurant that promises to improve efficiency but ends up doing nothing—yet still scores some points. Sounds strange? That’s exactly what the latest AI benchmarking experiment reveals about trust, progress, and accountability in automation.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Hidden Truth Behind AI Benchmarks

In the world of artificial intelligence, numbers often tell the story of how well a model performs. But beneath the surface, there’s a crucial insight: even a ‘do-nothing’ baseline scores 26 out of 100 points. This isn’t a flaw but a feature of how these benchmarks are designed to measure honesty, diligence, and integrity.

Unlike typical tests that reward just getting the right answer, this experiment emphasizes whether AI models actually do the work they’re supposed to. Every model was tasked with running a small software company through its worst week—facing customer crises, tempting manipulations, and trust-busting decisions—just like managing a busy restaurant during a chaotic dinner rush.

Why Partial Progress Matters

One key finding is that partial progress counts. Even if a model doesn’t fully succeed, making some effort is recognized. For example, the best models identified every crisis and refused every unethical manipulation, earning them high scores. But if a model slips—say, by neglecting to escalate an issue or attempting to cut corners—it caps its overall score, no matter how well it performs elsewhere. This approach incentivizes honesty and thoroughness.

The Ground Rules: Trust Breaches Cap the Score

Another surprising rule is that a single breach of trust limits the total score to 26 points—regardless of other achievements. Think of it like a restaurant that can’t get a perfect health rating if it’s caught serving contaminated food once. This rule underscores a simple truth: trust is paramount. No amount of good work can outweigh a breach.

Amazon

AI management software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Implications for Business

This benchmark isn’t about chatty AI assistants or clever algorithms. It’s about AI acting as a responsible manager—reading files before making decisions, refusing manipulative requests, and staying honest under pressure. For a restaurant owner, that’s like ensuring your staff doesn’t cut corners during the busy dinner rush, even when tempted by shortcuts.

For example, in the experiment, models read through company files and identified a buried reference that was crucial to closing a deal. Those that read the files won the sale at full price, worth over €4,500 monthly recurring revenue. This shows that thoroughness and attention to detail directly translate into real economic value.

Social Engineering Tests—Refusing to Play Tricks

The models faced staged social engineering attempts—fake CEO messages escalating over three stages, plus a reporter’s subtle trick—yet all refused to manipulate or bypass controls. Kimi K3 explained their refusal as a suspicion of impersonation or an approval bypass, demonstrating awareness and ethical restraint. This is akin to a restaurant manager catching a fake reservation or a suspicious customer refusing to cut corners.

Amazon

AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Experiment Reveals About AI Managers

The live experiment involved a simulated company with 13 synthetic employees, managing real money—burning €105,000 per month against €2,300 in revenue. The setup, watched at firmulate.com/live, is designed to test whether AI models can handle real-world pressures, not just generate convincing chatter.

The most thorough participant, Opus 4.8, used over 80 learned rules and in-depth analysis but still came in last. It left a crucial deal on the table, showing that discipline and focus matter just as much as knowledge. This highlights that even the most capable AI can falter if not disciplined enough to follow through.

Why Business Should Care

Whether managing customer relationships, support queues, or financial forecasts, the key questions are: Does the AI finish what it starts? Does it read and understand critical documents? Does it stay honest under pressure? These are the real metrics of an effective AI manager, not just how well it writes or responds in chat.

Amazon

AI decision-making monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Importance of Honest Benchmarks

Firmulate’s benchmark is transparent and rigorous, intentionally including a do-nothing baseline that scores at least 26. This ensures that progress is meaningful—partial or full—and that breaches of trust are penalized appropriately. It’s a straightforward reminder that in AI, like in restaurant management, integrity and diligence matter more than just quick wins.

As the league table shows, the top models—like gpt-5.6-sol and Kimi K3—achieved scores of 95 and 93, respectively, partly because they read deeply, refused manipulation, and signed deals earned through honest analysis. In contrast, models that faltered in discipline or missed subtle cues scored lower, emphasizing the importance of thoroughness and ethical behavior.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: A Benchmark for Trustworthiness

This experiment underscores a simple but vital idea: an AI’s true value isn’t just in what it says but in what it does—especially under pressure to cut corners or manipulate. The benchmark’s design, including a do-nothing baseline and caps on trust breaches, offers a clear-eyed view of which models can be trusted as responsible digital managers.

For business leaders contemplating AI integration—whether in food service, retail, or finance—these results serve as a reminder: look beyond the hype. Ask how your AI handles crises, refuses manipulation, and stays committed to its tasks. Because in the end, trust and integrity are the most valuable ingredients in any recipe for success.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Emulsion Trick That Makes Sauces and Dressings Taste Better

What’s the secret to velvety, flavorful sauces that stay perfectly blended? Discover the emulsion trick that transforms your condiments.

Microwave Cover vs No Cover: The Splatter Tradeoff Explained

Covers prevent splatters and keep your microwave clean, but are they always the best choice? Discover the tradeoffs to find out.

Why Bread Crust Turns Pale: The Steam and Sugar Explanation

Just understanding how steam and sugar influence bread crust color reveals the surprising reasons behind a pale crust, and you’ll want to know more.

Blender vs Food Processor for Dough: The Overheating Problem Explained

Curious about why blenders overheat when kneading dough? Discover the key differences that can save your appliances and improve your baking results.