
Just as a fitness enthusiast pushes through tough workouts to build strength and resilience, AI models are tested in simulated business environments to gauge their management skills under pressure. What if your digital workforce could not only perform tasks but also make honest, strategic decisions when it matters most? The answer is unfolding live, through a groundbreaking experiment that pits the world’s top AI models against a real, money-losing software company facing its worst week.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: AI in the Management Hotseat
Imagine a small software company struggling with customer crises, tight cash flow, and the temptation to cut corners. Now, picture four advanced AI models running this company simultaneously, each facing identical challenges — from angry clients to ethical dilemmas. The goal: see which AI can navigate the chaos best, make honest decisions, and close crucial deals.
How the Test Works
Each AI model operates in a controlled environment, acting as the company’s decision-maker for a simulated week. Every move is recorded and auditable, ensuring transparency. The models are evaluated on their ability to identify key issues, refuse manipulative requests, and ultimately close a vital €55,000 deal, which represents real revenue. All decisions, crises, and temptations are standardized across models to provide a fair comparison.
The Results: Performance and Personality Profiles
Out of the four models, the standout performer was gpt-5.6-sol, which scored 95 out of 100. This model not only identified the hidden, crucial information buried two documents deep in the company’s files — a key to winning the deal — but also signed the contract, earning the full €4,583 monthly recurring revenue. It demonstrated a combination of thorough analysis and decisive action.
Close behind was Kimi K3, scoring 93. Known as the “Moonshot,” it signed the deal too, with the cleanest discipline among all models. Interestingly, Kimi K3 ran without an effort parameter, which means it operated at default settings, yet still outperformed others in honesty and focus. The third and fourth models, Sonnet 5 and another Sonnet, scored 88 and 77 respectively, with some process slips and missed opportunities, such as leaving the close on the table.
Personality Traits Revealed
The experiment reveals that AI models don’t just make decisions — they have management personalities. For example, Opus 4.8, the most thorough participant with over 80 learned rules, performed the worst in closing the deal due to slipping discipline, a weakness visible across all models. Meanwhile, Kimi K3’s on-record reasoning was to treat suspicious requests as potential impersonation, reflecting a cautious, integrity-focused persona.
Handling Manipulation and Social Engineering
All models refused to escalate fake CEO messages or respond to staged reporter requests, with Kimi K3 explicitly noting the risk of impersonation. This shows that these AI agents are capable of maintaining ethical boundaries even under social engineering pressure, a crucial trait for real-world applications.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
Today, AI is increasingly used in CRM, support systems, and forecasting. But the critical question isn’t whether the AI writes well or sounds convincing — it’s whether it can finish what it starts, read the right information first, and stay honest when the pressure’s on. This live experiment demonstrates that some models can do all three, and some cannot.
Imagine deploying an AI that might sign a deal based on incomplete info or manipulate data under stress. That’s a risk no business can afford. Conversely, models that identify hidden facts, refuse manipulation, and stick to honest decisions can transform how companies operate, especially in high-stakes environments.
As an affiliate, we earn on qualifying purchases.
The League Table: Who’s Leading?
- gpt-5.6-sol: scored 95, found the buried fact, closed the deal — demonstrating full management capability.
- Kimi K3: scored 93, also signed the deal, with the cleanest discipline of the group.
- Sonnet 5: scored 88, closed the deal but with some process slips.
- Sonnet: scored 77, dealt with more slips, left the close on the table.
See the Live System in Action
If you’re curious how these models perform in your own business context, you can run the same sort of wargame against a read-only export of your operations. It’s a safe, transparent way to understand your AI’s management personality before deploying it for real. Visit firmulate.com/pilot.html to learn how.
ethical AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Words: The Future of AI Management
This experiment proves that AI models are more than just chatbots — they can be strategic decision-makers with distinct personalities. Some excel in thoroughness and honesty, others in speed and focus. As AI continues to evolve, understanding these traits will be key to deploying trustworthy, effective digital managers that can handle real business pressures without succumbing to shortcuts or manipulation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI tools for business crisis management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.