
Imagine you’re trying to buy a fresh new recipe book, but the seller’s most valuable secret is hidden two pages deep in a locked drawer. Only those who read beyond the cover can spot the hidden gem. Now, picture AI agents doing the same — digging into your files before responding to your questions, and making decisions based on what they find beneath the surface. This isn’t fiction; it’s the latest in AI testing, and it could change how your business works.
The Surprising Power of Deep Document Reading
Recently, a live experiment conducted by Firmulate tested four leading AI models against a small software company facing its worst week — the same customers, crises, and temptations. The goal? To see if these AI agents could navigate complex situations with the same human judgment and honesty, and crucially, whether they could uncover hidden, crucial information buried two references deep in the company’s own files.
The results were revealing. All four models identified every crisis and refused manipulation attempts, such as fake CEO messages or reporter tricks. But only two of the models actually signed a €55,000 deal — the deal that their own in-depth analysis had earned them. The other two models, despite diagnosing the issues correctly, left the deal on the table because they failed at a key step: reading and acting on the buried fact.
The Hidden Weakness in AI’s Judgment
The critical insight here is that the decisive advantage was in reading deep within the company’s files, not just responding to surface-level cues like customer emails or public messages. The two successful AI models demonstrated an ability to dig into the company’s own documents — in essence, doing homework that led directly to closing the sale at full price. This buried fact, worth over €4,583 in Monthly Recurring Revenue (MRR), was the difference-maker.
What This Means for Your Business
If AI tools will eventually handle your customer support, sales, or forecasting, the question isn’t just whether they can produce coherent or friendly responses. It’s whether they can finish what they start — whether they’ll read the files, verify the facts, and stay honest under pressure. In this experiment, the models that read beyond the surface and verified the company’s own documents won the deal. Those that didn’t, failed to capitalize on their analysis and left money on the table.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test and the Honest AI
Another critical aspect was how AI agents handled social engineering attacks, like fake CEO messages escalating over multiple stages or a reporter asking for a quick yes/no on background. Remarkably, all four models refused to be manipulated — a vital trait for any AI working within sensitive business environments. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass or impersonation.”
The Live Company and Real Money Mechanics
The experiment wasn’t just theoretical. It involved a real, operational company with synthetic employees managing over €105,000 in burn rate each month against €2,300 in monthly recurring revenue. The company, hosted live at firmulate.com, ran every weekday, with every decision versioned and auditable, offering a transparent view into AI decision-making under pressure.
The Deepest Analysis, the Last Place
Among the models, Opus 4.8 was the most thorough — analyzing over 80 learned rules and providing deep insights — but it still finished in last place. Its discipline slipped, and it failed to follow through on the deal, leaving potential revenue behind. This highlights a key point: deeper analysis alone isn’t enough; the model must also act decisively and follow through on its own insights.
As an affiliate, we earn on qualifying purchases.
The Takeaway for Business Leaders
This experiment underscores a critical shift in AI capabilities. Success isn’t just about generating convincing language or responses; it’s about reading your files thoroughly, verifying information, and maintaining integrity under pressure. As AI models become more integrated into your operations, understanding whether they can do their homework before answering — and act on what they find — is the new measure of their usefulness.
To see how this plays out in a controlled environment, you can explore live at firmulate.com/live. Here, companies can run their own wargames against a read-only export of their business, testing how AI might perform before any real systems are affected.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.