
Imagine a restaurant struggling with customer complaints and staff shortages. Could an AI step in, make tough decisions, and steer the business through chaos? It turns out, AI models are not just good at chatting—they are beginning to manage real companies under real pressure. But do they do it well enough to trust? That’s what a groundbreaking experiment by Firmulate reveals.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: Putting AI to the Test in a Live Business
In a daring experiment, four state-of-the-art AI models were tasked with running a small software company during its worst week—crises included. Every decision they made was recorded, auditable, and identical across models, ensuring a fair comparison. The goal? To see if these models could manage crises, avoid manipulation, and close a crucial €55,000 deal at the end of the week.
The Models and Their Scores
According to the final standings, the models showed remarkable abilities:
- GPT-5.6-SOL scored 95 points, found hidden information in company documents, and secured the deal.
- Kimi K3 scored 93, signed the deal without effort parameters, and maintained the cleanest discipline.
- Sonnet 5 scored 88, also closed the deal but with some process slips.
- Fable 5 scored 77, closed the deal too, yet made more mistakes along the way.
Interestingly, a baseline score of 26 showed how much progress these models made, with full performance achieved only by the top two.
As an affiliate, we earn on qualifying purchases.
Decisiveness and Ethics Under Pressure
In the face of managed crises, the models demonstrated integrity: all four identified every crisis, refused every attempt at manipulation, and upheld honesty even when fake CEO messages escalated over multiple stages. Kimi K3 justified its decisions by treating suspicious requests as potential impersonations, showing a cautious and ethical approach.
What Made the Difference?
The key to winning the deal wasn’t just about diagnosing problems but digging deeper into the company’s own files. The models that examined references buried two documents deep in the company’s files uncovered critical insights that others missed. This extra step allowed them to identify hidden issues, leading to successful negotiations and full-price deals—worth over €4,500 in recurring monthly revenue (MRR).
As an affiliate, we earn on qualifying purchases.
The Real Company Behind the Experiment
The experiment unfolds within a real, functioning software firm with 13 synthetic employees and actual money mechanics. The company burns €105,000 each month against a modest €2,300 MRR, with a public cash countdown—a stark reminder of the pressure to perform. Every workday, the decision-making process is versioned, transparent, and observable, making this a rare window into AI’s practical capabilities in management.
The Profiles and Lessons
Among the models, OPUS 4.8 stood out as the most detailed and thorough—analyzing over 80 learned rules and conducting deep assessments. Yet, it left some opportunities unseized, like leaving a close deal on the table or slipping into a locked department instead of escalating issues properly.
By contrast, Kimi K3 operated without an effort parameter, running at default settings, but still managed to close the deal cleanly. This indicates that different models exhibit distinct managerial personalities—some meticulous, others more straightforward—and all are measurable in their decision styles.
AI decision-making platforms for enterprises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
This experiment underscores an essential point: the question with AI in management is no longer just about how well it writes or chats. It’s whether the AI can finish what it starts, understand vital documents, stay honest under pressure, and ultimately deliver value. As businesses increasingly deploy AI agents in CRM, support, or forecasting, these qualities become critical.
Try It Yourself
If you’re curious about how your own enterprise might fare against AI decision-makers, you can run similar tests in a safe, read-only environment. These simulations expose the AI’s decision patterns, strengths, and weaknesses—giving you a clearer picture before you hire or integrate AI into your operations. Visit firmulate.com/quiz.html to test your company’s AI management profile.
AI negotiation and deal closing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Takeaway
In a world where AI increasingly touches every aspect of business—from customer service to strategic decision-making—the ability to trust your AI managers is vital. This experiment shows that some models excel at not only identifying crises but also resisting manipulation and closing deals at full value. The future of management may very well depend on understanding these AI personalities and choosing the right one for your company’s unique needs.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
