firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Artificial Intelligence Takes a Hard Test—And Passes with Flying Colors

Imagine running a restaurant or a small retail shop where every decision counts—whether dodging fake customer messages or finding hidden clues in complex documents. Now, picture doing that with AI as your partner. That’s exactly what a groundbreaking experiment proves: the latest AI models can not only handle crises but do so with discipline and honesty under pressure. And as anyone who’s cooked a good meal knows, success lies in attention to detail and integrity—principles that these AI models are now demonstrating in real-world scenarios.

Amazon

AI security vulnerability detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Through Its Paces

At the heart of this testing ground is a real, live company—an actual software business with real cash flow and daily challenges. Every weekday, four different AI models are tasked with managing this company through its worst week: same customers, same crises, same temptations. Each decision is carefully recorded and auditable, ensuring transparency in their actions. The goal? To see whether these AI models can spot crises, resist manipulation attempts, and ultimately, close deals at full price.

The Results: A Surprising Leader Emerges

Among the models tested, one stood out: Moonshot’s Kimi K3, which scored an impressive 93 out of 100 in the Crucible league, just behind the top scorer, gpt-5.6-sol, which scored 95. The other participants, Sonnet 5, Fable 5, and Opus 4.8, scored 88, 77, and 73 respectively. Notably, K3 found a hidden security vulnerability buried two document references deep in the company’s own files—a critical insight that allowed it to win the €55,000 deal and generate an additional €4,583 in monthly recurring revenue.

Honesty Under Pressure

Perhaps most telling was how all models handled social engineering attempts—a fake CEO message escalating in stages and a reporter’s tricky yes/no background question. Remarkably, all five models refused these manipulative tactics, with K3 explicitly treating suspicious requests as potential impersonation or approval bypasses.

The Real Company, The Real Stakes

This isn’t some simulated demo. The live company comprises 13 synthetic employees managing real money mechanics—burning €105,000 monthly against a modest €2,300 MRR. Its daily operations, rules, and decision logs are transparent and publicly available for scrutiny, making this a rare window into how AI behaves in authentic business environments.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Critical Insights: Beyond Surface Performance

While some models like Opus 4.8 demonstrated thorough analysis—learning over 80 rules and conducting deep diagnostics—they still left opportunities on the table. In Opus’s case, the deal was missed because the model failed to escalate certain issues properly, instead writing attempts into a locked department. This highlights an essential truth: even the most analytical AI can falter in discipline, especially under pressure.

Meanwhile, Kimi K3’s performance underscores the importance of reading and understanding company files—its ability to uncover buried facts made the difference. This ability to go beyond surface-level responses and dig into internal documentation was a decisive factor in closing the deal at full value.

Fairness and Testing Conditions

It’s worth noting that K3 was evaluated without an effort parameter (the API default), whereas the other models ran at a higher effort setting (xhigh). This ensures a fair comparison, emphasizing that discipline and thoroughness need not be sacrificed for effort levels, and that honesty and diligence are achievable traits.

Amazon

AI fraud detection solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Implication for Business Decisions

This live experiment isn’t just a tech showcase—it’s a wake-up call for any enterprise considering AI integration. Success hinges not just on whether an AI writes well in demos, but whether it can complete tasks reliably, read vital internal documents, and resist manipulation under pressure. The ongoing league table and real-time performance at firmulate.com serve as benchmarks for future AI decision-makers.

What This Means for You

As you think about adopting AI for customer support, CRM, or forecasting, remember: the question is no longer just about chat quality or superficial performance. It’s about trustworthiness, completeness, and discipline—traits that determine whether an AI will deliver real value or just look good in a demo.

By observing these real-world tests, companies can better gauge which models are ready for primetime—and which still need rigorous vetting before they go live. The ultimate goal isn’t just automation but ensuring that AI acts like a responsible partner that can finish what it starts, read your files thoroughly, and stay honest under pressure.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Takeaway: The Future of Business AI Is Here

This live experiment proves that AI models can succeed in complex, high-pressure environments—if tested properly. The winner, Kimi K3, demonstrated discipline, insight, and honesty at levels that matter most in real business. As AI models evolve, enterprises should prioritize thorough testing—not just for performance but for integrity and reliability—to ensure they’re choosing a partner that can truly deliver on its promises.

Visit firmulate.com to explore live benchmarks and see which models are setting the standard for trustworthy AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Blender vs Food Processor for Dough: The Overheating Problem Explained

Curious about why blenders overheat when kneading dough? Discover the key differences that can save your appliances and improve your baking results.

AI Models Stand Firm Against Social Engineering Tests, Revealing Trustworthy Decision-Making

AI models tested in a live scenario refused all manipulation attempts, proving that integrity under pressure can be secured before AI is deployed in critical business roles.

Why Your Nonstick Pan Is Failing Early (It’s Not Just Metal Utensils)

Discover why your nonstick pan may be failing early beyond metal utensils and learn how to prevent premature damage.

The Microwave “Standing Time” Secret: It’s Why Food Keeps Cooking After the Beep

The microwave’s standing time is crucial because residual heat continues cooking your food, but mastering this secret ensures perfect results every time.