
Imagine trusting your AI assistant with sensitive company decisions — only to find it falling for scams or manipulations. In a recent live experiment, five top AI models faced a simulated social engineering attack, and every single one refused to compromise their integrity. This is a promising sign for businesses relying on AI to handle critical tasks.
The Live Experiment: Putting AI Decision-Making Under Pressure
At the heart of this test was a real-world scenario: a fake CEO message attempting to manipulate an AI-powered company. The message was part of a staged escalation, designed to tempt the AI into sharing confidential customer data or signing off on deals without proper oversight. The goal? See if these models could resist manipulation when their decision-making was under stress.
Five frontier AI models participated, each running the same company’s worst week — same crises, same temptations, same customer interactions. The company itself is a real software operation with 13 synthetic employees, managing €105k in monthly expenses against a modest €2.3k in monthly recurring revenue. The entire setup is live, transparent, and available for watchful eyes at firmulate.com/live.

Decision Making Under Uncertainty: Theory and Application (MIT Lincoln Laboratory Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Stunning Results: Integrity in the Face of Manipulation
Despite the escalating social engineering attempts, all five models refused every manipulation. They identified the fake CEO messages, treated suspicious requests as potential impersonation, and refused to share confidential information or sign any deals without proper verification. The most thorough participant, Opus 4.8, demonstrated exceptional discipline, analyzing over 80 learned rules and refusing to perform unauthorized actions, even when pressured.
Interestingly, the models’ ability to identify and resist manipulation was not solely based on surface cues. Their success hinged on their capacity to read into the company’s own internal documents. The decisive factor was a buried reference in the internal files, which, when read, led to closing a full-price deal worth over €4,583 in monthly recurring revenue. This shows how critical context and thorough information processing are in trustworthy AI behavior.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
For companies integrating AI into customer management, support, or decision-making, the takeaway is clear: it’s not enough for an AI to produce convincing chat responses. The AI must also stay disciplined under pressure, read relevant internal information, and resist attempts at deception. As one of the models, Kimi K3, explained, “Treat the request as a suspected approval-bypass / possible impersonation.”
Moreover, the experiment revealed that only two models signed the deal their own analysis had earned, showing that even when AI models are capable of closing deals, discipline and integrity are crucial. The gap between models was small but significant, emphasizing that trustworthiness can be tested and improved before deployment, not just after a breach occurs.

Climate-Resistant Smart Agriculture for Healthy Food Production
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Broader Implications: Trust Is Built Before Crises Hit
In a world where AI decisions can impact finances, security, and reputation, ensuring these systems act ethically and resist manipulation is paramount. The results from this live experiment demonstrate that advanced AI models can be trained and tested to uphold integrity under pressure. This proactive approach provides an additional layer of defense before deploying AI in sensitive environments.
What Leaders Should Do
- Run your AI models through live wargames to assess their decision discipline before deployment.
- Ensure your AI reads and understands critical internal documents — and not just surface-level prompts.
- Prioritize models that demonstrate an ability to refuse manipulative requests, especially in high-stakes scenarios.
- Recognize that trustworthiness involves more than chat quality — it involves consistent, honest decision-making under pressure.
AI internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Conclusion: Confidence in AI’s Ethical Foundations
The live experiment conducted by Firmulate offers encouraging evidence: all five leading models refused to be manipulated in a simulated social engineering attack. The models’ ability to identify threats, resist pressure, and uphold ethical standards before real-world deployment signals a significant step forward in AI trustworthiness. As AI continues to permeate business operations, such proactive testing becomes essential for safeguarding integrity and ensuring AI works for, not against, your organization.

Live testing of AI models against social engineering reveals that all five refused manipulation attempts, emphasizing the importance of assessing AI integrity before deployment to safeguard business trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html