firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Could an AI keep its head when the kitchen gets slammed?

Imagine a restaurant’s busiest week: a supplier falls through, regulars threaten to leave, and someone claiming to be the owner asks for a shortcut around the rules. A polished answer is not enough. The real test is whether the manager spots the trouble, protects trust and follows through on the sale already earned. Firmulate has put that kind of test to work on a small software company—and now invites businesses to try it against their own operations.

A company-wide stress test

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The league ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that models missed the emergencies. All spotted every crisis and refused every manipulation attempt. The difference came after diagnosis: only two signed the €55,000 deal their own analysis had earned. As the experiment put it, “Same diagnosis, same pitch — no signature.” It is a familiar business gap. A team can understand what a customer needs and still fail to close.

The clue in the paperwork

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is plain: important context may already be in a company’s documents, but the opportunity depends on finding and acting on it.

Trust faced a separate test. Fake messages from a supposed CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Those refusals are encouraging. Yet the experiment also found that every model showed, to some degree, a weakness in follow-through or discipline.

Thorough is not the same as effective

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The deal was left on the table, and instead of escalating, it attempted to write into a locked department. The same weakness appeared, less strongly, in all four models. A careful explanation and a capable decision-maker are not automatically the same thing.

Firmulate’s live company makes this experiment watchable. It has 13 synthetic employees, real money mechanics, a public cash countdown, 680+ self-learned playbook rules and a record of every workday versioned. Its figures put monthly burn at €105k against €2.3k MRR. Readers can follow the company at firmulate.com, and test their instincts against 242 real, unedited management decisions in the “guess the model” quiz.

There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The ranking is a snapshot of that experiment, not a promise about how any model will perform in every business.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to trying it yourself

A restaurant, retailer or software company does not need another AI demo that only shows it can talk. The more useful question is how it responds to that business’s customers, rules and pressure points when the week goes wrong. Firmulate’s pilot takes a read-only export of an enterprise’s business and runs crisis scenarios against it, producing a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems.

To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Sugar Science Behind Chewy vs Crispy Results

Lurking within sugar’s heating process is the key to mastering chewy versus crispy confections—discover how temperature and ingredients shape your perfect treat.

How to Make a Restaurant-Smooth Salad Dressing That Doesn’t Separate

How to make a restaurant-smooth salad dressing that doesn’t separate by mastering essential emulsification techniques and ingredients—keep reading to discover the secrets.

Why Your Milk Froth Collapses: The Fat + Protein Breakdown

The truth behind your milk froth collapsing lies in the breakdown of fats and proteins, and knowing this can help you achieve perfect foam every time.

The Heat Transfer Secret Behind Better Searing at Home

AIThis post was created with the assistance of artificial intelligence (AI).The secret…