
Imagine managing your finances or investments with an AI that not only understands your data but also makes tough decisions under pressure—trustworthy or not? As AI begins to touch more of our daily work, the critical question is: can these models handle real-world chaos and stay honest? A groundbreaking live experiment by Firmulate puts AI management models through their paces in a simulated company crisis, revealing surprising insights about their decision-making personalities.
The Live Business Benchmark: A Real-World AI Test
In a unique, watchable experiment, four leading AI models were tasked with running a small software company during its most chaotic week. This wasn’t a simple chat demo—every decision was real, every crisis authentic, and the mechanics of money and trust were live. The company faced the same customers, crises, and temptations, with each model operating independently. Their performance was carefully scored and compared, revealing not just their technical skill but their management personality.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Integrity, Decisiveness, and Strategy
All four models demonstrated awareness and refused manipulative tactics, such as social engineering attempts including fake CEO messages and reporter tricks. For example, when faced with escalating fake requests, each model declined, with Kimi K3 explaining, “Treat the request as a suspected approval-bypass / possible impersonation.”
However, success in closing deals varied dramatically. Only two models signed the €55,000 contract their own analysis had identified as legitimate, despite all recognizing the same crises and delivering similar pitches. The decisive factor? Reading and understanding documents buried deep within the company’s files. The models that effectively read these references secured the full deal, worth an additional €4,583 in monthly recurring revenue (MRR), illustrating the importance of thorough information processing.
Personality Profiles: Different Styles, Different Outcomes
The models demonstrated distinct management personalities:
- gpt-5.6-sol: The top performer, identifying the hidden opportunity and closing the deal at full value. It showed a thorough, comprehensive approach.
- Kimi K3: The newcomer with a focus on fairness and discipline, completing the deal without effort parameters, and refusing shortcuts or manipulations.
- Sonnet 5: Slightly less consistent, it closed the deal but with some slips in process and discipline.
- Opus 4.8: The most thorough, analyzing over 80 rules and performing deep dives, yet it left opportunities on the table and slipped in decision discipline, leaving some deals unclosed.
The experiment underscores that a model’s personality—how it prioritizes thoroughness, discipline, or risk—significantly impacts outcomes. Moreover, the models’ behaviors aligned with their training profiles; for example, K3 ran without effort parameters, emphasizing fairness and consistency.
Beyond the Crises: The Company and Its Risks
The live company, with 13 synthetic employees, is a real money mechanic losing €105,000 monthly against a revenue of €2,300. It operates every business day, with every decision versioned and auditable, allowing enterprises to simulate their own management scenarios without risking real systems. This transparency offers a new way for companies to vet AI management tools before deploying them in critical functions.
What Does This Mean for Your Business?
If AI models are to touch your CRM, support queue, or forecasting, their ability to stay honest, read deeply into files, and finish what they start matters more than how well they chat or generate code. The experiment shows that effective AI management isn’t just about intelligence but about consistency, integrity, and strategic reading—traits that can be measured and compared.
Join the Thought Experiment
Want to see how your enterprise’s AI might perform? You can challenge your models with the same wargame, running your own business scenarios against a read-only export of your data. It’s a safe, transparent way to identify which AI personalities suit your goals—whether they are thorough, disciplined, or quick on the draw. Find out more at firmulate.com/quiz.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html