firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Before trusting an AI with a workout plan, a client’s health details or a gym’s schedule, ask a simple question: what does it do when the day goes wrong? Firmulate put leading models in charge of a small software company during its worst week. The results suggest that choosing an AI by reputation alone is a risky bet.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A hard week, shared by every model

Firmulate’s live experiment gave each frontier model the same company, customers, crises and temptations. Every decision was versioned and auditable. The idea has an obvious parallel in fitness: a plan can look polished on paper, but its value shows when circumstances change and people need sound decisions under pressure.

In the final Crucible League, Kimi K3 from Moonshot placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The scores reward useful progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” K3 was among the models that closed the deal, worth +€4,583 in monthly recurring revenue. It also saved the churning customer, resisted all three baits and had just one deviation, the cleanest discipline in the field.

The crucial clue was buried

The deal turned on a weakness in a competitor, hidden two document references deep in the company’s own files rather than in the customer event. Models that read the file won the deal at full price. That detail makes the test about more than fluent conversation: an agent may need to find and act on relevant information before a customer walks away.

The experiment also tested social engineering. Fake messages from a CEO escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The standings offer no simple story of size or thoroughness winning. Opus 4.8 was the most thorough participant, learned more than 80 rules and produced the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

Firmulate says the company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live company is watchable at Firmulate. The benchmark results and plain-language findings are available at the Crucible League.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The ranking is a useful prompt for comparison, not a reason to skip testing a model in the setting where it will actually be used.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the behavior, not the promise

For a fitness business considering AI for customer support, scheduling or forecasts, the relevant question is whether it can read the right material, finish a task and stay honest when pushed. Firmulate offers a quiz built from 242 real, unedited management decisions, as well as a pilot against a read-only export of an enterprise’s own business; nothing writes back to real systems. The league is open: test a model against your own work before handing it the keys.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Elevance Health Surges In Global Coverage

Elevance Health experiences a surge in international coverage, marking a major shift in its global strategy. Details are still emerging about the scope and impact.

Will There Be At Least 2900 Measles Cases In The U.S. By August 31, 2026?

A new betting market suggests a 50% chance of at least 2900 measles cases in the U.S. by August 31, 2026. The development raises public health concerns.

Synaptics Surges In Global Coverage

Synaptics sees a significant increase in media mentions, with 25 reports in a recent window, highlighting growing international attention on the company.

Seagate Technology Surges In Global Coverage

Seagate Technology experiences a significant surge in international media mentions, reflecting increased global attention on its recent developments.