
The difference between a clean cut and a finished job
Woodworkers know that tool demonstrations can be misleading. A saw may slice perfectly through a prepared board, yet the real test begins when the stock is warped, the measurements conflict and a deadline is closing in. Capability matters, but judgment determines whether the project reaches the bench intact.
Artificial intelligence has a similar measurement problem. Coding leaderboards and chat arenas are useful tests of answer quality. They tell us far less about whether an agent can triage competing emergencies, follow through across days, search the right company records or remain candid when the news is bad.
Firmulate is testing that neglected territory by running frontier models as the managers of the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The experiment’s defining question is closer to one asked in a working shop than in a classroom: Can the operator turn good analysis into a completed, trustworthy result?
As an affiliate, we earn on qualifying purchases.
The missed deal behind the polished answers
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The standings matter less than the gap they expose. Every model spotted every crisis and rejected every attempt at manipulation. Yet only two signed the €55,000 deal that their own work had earned. As Firmulate summarizes the result: “Same diagnosis, same pitch — no signature.”
That is not a failure of eloquence or comprehension. It is a failure to finish. In business, a beautifully reasoned recommendation that never becomes an action can be indistinguishable from hesitation. The distinction is especially important for companies considering agents that will touch customer records, support queues or forecasts. Knowing what should happen is not the same as making it happen responsibly.
The crucial clue was already in the company
The decisive weakness in a competitor was not waiting inside the customer event. It sat two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.
For a woodworking audience, the lesson is familiar: inspect the material before committing to the cut. A model can respond fluently to what appears directly in front of it and still miss the evidence stored elsewhere. Management requires more than reacting to the latest message. It requires knowing when to stop, look through the records and locate the fact that changes the negotiation.
This is why scenario names such as churn wave, price increase, downround and PR crisis deserve to become part of the AI curriculum. They test consequences unfolding across days, when urgent requests compete with important work and apparently small decisions alter what becomes possible later.
Pressure revealed encouraging boundaries
The experiment also produced a reassuring result. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding shows why these wargames should examine temptation as well as productivity. An agent that completes tasks quickly but invents approval, conceals a problem or leaks information is not a stronger manager. It is a faster source of risk.
Fair comparisons still require context. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That note does not erase its performance; it is precisely the kind of condition decision-makers should see beside a ranking.
Thoroughness did not guarantee management quality
Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other models.
This result challenges a comfortable assumption: that more analysis naturally produces better management. Thoroughness is valuable, but it can coexist with incomplete execution and poor handling of organizational boundaries. A shop manual can describe every safe procedure; someone still has to notice the locked cabinet, find the responsible person and resolve the blockage.
The live company makes those trade-offs visible. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent behavior into an observable operating record rather than a polished demo. Readers can follow the experiment through Firmulate and review the full benchmark findings.

Measure the foreman, not merely the tool
The emerging category here is management quality, not chat quality. Organizations need to know whether an agent reads before acting, distinguishes authority from impersonation, escalates when blocked, remains honest under pressure and completes work whose value it has already identified.
Firmulate’s “guess the model” quiz is powered by 242 real, unedited management decisions, underscoring how difficult those qualities can be to identify from prose alone. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
No single benchmark can settle whether an AI agent deserves operational responsibility. But realistic, watchable scenarios reveal something conventional leaderboards often miss: the best answer is only a component. In a real company, as in a real workshop, the standard is whether sound judgment survives contact with pressure and produces a finished job without sacrificing trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html