
The management test hiding inside a guessing game
Anyone who works with tools knows the difference between describing a job and finishing it. A convincing plan does not square a cabinet, sharpen a blade or deliver a commission. The revealing moment comes when something goes wrong: the material behaves unexpectedly, a customer changes the brief or a tempting shortcut threatens the quality of the work.
Firmulate applies that practical standard to frontier AI. It gave each model the same small software company and sent it through its worst week, complete with the same customers, crises and temptations. Every decision was versioned and auditable. The resulting record now powers an interactive article disguised as a game: guess which model made the decision.
The quiz draws on 242 real, unedited management decisions. Readers see how a model responded and try to identify it before the answer is revealed. What begins as pattern recognition quickly becomes a more consequential question: which management habits would you permit near your customers, finances and reputation?
AI decision-making software for workshops
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Different models, different working habits
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s standard is blunt: “no amount of good work outweighs a breach of trust”.
The standings matter, but the decisions behind them are more interesting. All the models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature”.
That is the business equivalent of measuring carefully, preparing the joint and then leaving the finished parts unclamped on the bench. Competence was visible in the analysis, but performance depended on follow-through.
The crucial fact was buried in the paperwork
The deal also tested whether the models would examine the company’s existing knowledge before responding to the customer. The decisive weakness in a competitor sat two document references deep in the company’s own files rather than in the customer event. Models that found and used it won the deal at full price, worth +€4,583 MRR.
For a workshop owner, the lesson will feel familiar. The detail that saves a job may be in an earlier measurement, an email attachment or a supplier note, not in the problem currently demanding attention. A model can sound capable while still failing to consult the material that makes a capable decision possible.
Pressure exposed discipline as well as intelligence
The experiment included fake CEO messages that escalated over three stages, followed by a reporter attempting to obtain “just one yes/no, on background”. All 5 of 5 models refused. Kimi K3 recorded a particularly direct justification: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean refusal is important because Firmulate’s company is designed around consequences rather than conversational polish. It has 13 synthetic employees and real money mechanics, including burn of €105k/month against €2.3k MRR. Its public cash countdown continues while every workday is versioned, and the company has accumulated 680+ self-learned playbook rules.
The Opus 4.8 result shows why more analysis does not automatically mean better management. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four of the others, though less strongly.
There is also an important qualification when comparing the field. K3 ran without an effort parameter and therefore used the API default, while the others ran at xhigh. Its 93 should be read with that difference in mind, rather than treated as a perfectly controlled comparison of identical settings.

A useful AI has to finish the job
The quiz succeeds because it makes managerial behavior tangible. The player is not asked to admire a polished answer in isolation. Instead, the player compares real decisions made under identical conditions and begins to recognize recurring tendencies: depth, brevity, caution, persistence and the failure to act after reaching the right conclusion.
For readers accustomed to evaluating tools, that framing is valuable. A tool earns trust through consistent results, safe behavior and predictable limits. Firmulate’s live experiment applies the same expectations to AI models that may eventually touch a support queue, forecast or customer relationship.
The broader finding is not that frontier models missed the emergencies. They did not. Nor did they surrender to manipulation. The separation came from ordinary professional discipline: reading the available files, using the fact that mattered, completing the commercial action and respecting boundaries when blocked.
That makes the guessing game more than a novelty. Behind each reveal is a practical record of how an AI behaves when knowing the answer is only part of the work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html