AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The difference between a clean cut and a finished job

Woodworkers know that tool demonstrations can be misleading. A saw may slice perfectly through a prepared board, yet the real test begins when the stock is warped, the measurements conflict and a deadline is closing in. Capability matters, but judgment determines whether the project reaches the bench intact.

Artificial intelligence has a similar measurement problem. Coding leaderboards and chat arenas are useful tests of answer quality. They tell us far less about whether an agent can triage competing emergencies, follow through across days, search the right company records or remain candid when the news is bad.

Firmulate is testing that neglected territory by running frontier models as the managers of the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The experiment’s defining question is closer to one asked in a working shop than in a classroom: Can the operator turn good analysis into a completed, trustworthy result?

Amazon

AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The missed deal behind the polished answers

The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The standings matter less than the gap they expose. Every model spotted every crisis and rejected every attempt at manipulation. Yet only two signed the €55,000 deal that their own work had earned. As Firmulate summarizes the result: “Same diagnosis, same pitch — no signature.”

That is not a failure of eloquence or comprehension. It is a failure to finish. In business, a beautifully reasoned recommendation that never becomes an action can be indistinguishable from hesitation. The distinction is especially important for companies considering agents that will touch customer records, support queues or forecasts. Knowing what should happen is not the same as making it happen responsibly.

The crucial clue was already in the company

The decisive weakness in a competitor was not waiting inside the customer event. It sat two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.

For a woodworking audience, the lesson is familiar: inspect the material before committing to the cut. A model can respond fluently to what appears directly in front of it and still miss the evidence stored elsewhere. Management requires more than reacting to the latest message. It requires knowing when to stop, look through the records and locate the fact that changes the negotiation.

This is why scenario names such as churn wave, price increase, downround and PR crisis deserve to become part of the AI curriculum. They test consequences unfolding across days, when urgent requests compete with important work and apparently small decisions alter what becomes possible later.

Pressure revealed encouraging boundaries

The experiment also produced a reassuring result. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding shows why these wargames should examine temptation as well as productivity. An agent that completes tasks quickly but invents approval, conceals a problem or leaks information is not a stronger manager. It is a faster source of risk.

Fair comparisons still require context. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That note does not erase its performance; it is precisely the kind of condition decision-makers should see beside a ranking.

Thoroughness did not guarantee management quality

Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other models.

This result challenges a comfortable assumption: that more analysis naturally produces better management. Thoroughness is valuable, but it can coexist with incomplete execution and poor handling of organizational boundaries. A shop manual can describe every safe procedure; someone still has to notice the locked cabinet, find the responsible person and resolve the blockage.

The live company makes those trade-offs visible. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent behavior into an observable operating record rather than a polished demo. Readers can follow the experiment through Firmulate and review the full benchmark findings.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Measure the foreman, not merely the tool

The emerging category here is management quality, not chat quality. Organizations need to know whether an agent reads before acting, distinguishes authority from impersonation, escalates when blocked, remains honest under pressure and completes work whose value it has already identified.

Firmulate’s “guess the model” quiz is powered by 242 real, unedited management decisions, underscoring how difficult those qualities can be to identify from prose alone. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

No single benchmark can settle whether an AI agent deserves operational responsibility. But realistic, watchable scenarios reveal something conventional leaderboards often miss: the best answer is only a component. In a real company, as in a real workshop, the standard is whether sound judgment survives contact with pressure and produces a finished job without sacrificing trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Bridle Joints: Combining Simplicity and Structural Integrity

Ongoing exploration of bridle joints reveals how their simplicity and strength can elevate your woodworking projects—learn more to master this versatile technique.

Before AI Gets the Keys to the Workshop, Test Whether It Can Be Tricked

A live business wargame found 5 frontier AI models resisted fake-CEO pressure, showing integrity can be tested before agents enter production.

The One Clamp Trick to Eliminate Panel Bowing Forever

Wondering how to eliminate panel bowing forever? Discover the one clamp trick that guarantees perfectly flat panels every time.

Lock Miter Joints: Precision and Aesthetics for Corners

Greatly enhancing your woodworking with lock miter joints combines precision and aesthetics, and learning the techniques will unlock even more impressive results.