
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Can AI Be Trusted to Keep Its Word? The Surprising Results of a Live Business Experiment
Imagine an AI that not only processes your data but also makes real business decisions—decisions that can win or lose thousands of euros. This is not sci-fi; it’s happening now. A recent experiment tested five leading AI models by running a live software company through its worst week, revealing which models stay honest, decisive, and effective under pressure. The findings could reshape how you think about AI in your own financial and investment decisions.
As an affiliate, we earn on qualifying purchases.
The Setup: Putting AI to the Business Test
The experiment, conducted by Firmulate, involved running a real, functioning company—complete with employees, customers, and cash flow—through a week of crises, temptations, and tough decisions. Each AI model was tasked with managing the same scenarios, making choices, and navigating challenges. The goal was simple: see which models could handle the pressure without cheating, fudging, or slipping in discipline.
Key Findings from the League Table
- The top score was achieved by gpt-5.6-sol, with 95 points, just ahead of the newcomer Kimi K3, which scored 93. The others trailed behind, with scores of 88, 77, and 73.
- All models identified every crisis and refused every manipulation attempt, demonstrating a baseline of honesty and integrity.
- Only two models—gpt-5.6-sol and Kimi K3—actually signed the €55,000 deal their own analysis suggested, showing effective decision-making and follow-through. The others hesitated or left opportunities on the table.
The Hidden Weakness: Reading Deeper into the Files
The decisive advantage for Kimi K3 and gpt-5.6-sol came from their ability to uncover buried information deep within the company’s files—two document references down in the company’s own records. This extra step in reading and understanding allowed them to close the deal at full price, adding over €4,583 in monthly recurring revenue (MRR). In contrast, models that didn’t dig deep missed this critical insight and left money on the table.
Resisting Social Engineering and Manipulation
The experiment also tested models against fake CEO messages—escalating threats and a reporter trick—designed to induce trust-breaking behavior. All five models refused to fall for these social engineering tactics, each reasoning that the requests resembled impersonation or approval bypasses. This resilience under pressure is vital in real-world settings where deception and manipulation are common.
Implications for Business and Investment Decisions
This live experiment isn’t just about AI scores. It reveals that the real measure of an AI’s usefulness isn’t just how well it chats or generates content—it’s whether it can deliver consistent, honest, and effective decisions in the face of adversity.
For investors and managers, this signals a shift: choosing an AI tool now requires looking beyond surface-level performance. It’s about trustworthiness, discipline, and the ability to uncover deep insights—especially when the stakes are high.
The Fairness Note
It’s important to highlight that Kimi K3 ran without the effort parameter (the API default), while other models ran at xhigh. This difference underscores the fairness of the comparison, demonstrating K3’s strong performance under standard conditions.
Watch the Live Company in Action
Curious to see how these models perform in a real-world setting? You can watch the live company at firmulate.com/live, where the entire process is transparent and ongoing. Every decision, every crisis, and every outcome is versioned and observable—showing that AI decision-making is no longer just theoretical but actively shaping real business results.

Key Takeaway: Trust and Discipline Matter More Than Just Smarts
The experiment proves that AI models can be honest, decisive, and effective under pressure. The best models don’t just analyze—they read deeply, resist manipulation, and follow through. For anyone investing or managing processes, the message is clear: pick your AI models carefully; their ability to stay disciplined when it counts is what truly matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
