
Would Your Apprentice Read the Spec Sheet — or Just Eyeball It?
Every woodworker knows the rule: measure twice, cut once. The joiner who skims the drawing, guesses at the mortise depth, and “knows” the board is straight ends up buying a new plank. The one who flips back two pages to the fine print — where the buried detail about the material’s tolerance lives — gets it right the first time, at full price.
It turns out AI models running real businesses suffer from exactly the same split. In a live, watchable experiment by Firmulate, four frontier AI models were each handed the same small software company during its worst week. All of them spotted every crisis. All of them refused every attempt to con them. But only some of them flipped back the two pages — and that single habit decided a €55,000 deal.
manual reference checker for documents
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Wargame
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the final July 2026 “Crucible” league, each model steered the same firm through identical customers, identical disasters, and identical invitations to cut corners. Every decision was versioned and auditable, like a cut list you can check afterwards.
The final standings: gpt-5.6-sol won with 95 points, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, doing nothing at all still scored 26 — partial progress counts — but a single breach of trust caps the whole total. As the rules put it, “no amount of good work outweighs a breach of trust.”
The Fact Buried Two Documents Deep
Here’s the finding that matters. Somewhere in the company’s own files — not in the customer conversation, not in the dramatic moment — sat a decisive competitor weakness. But it wasn’t on page one. It sat two document references deep, the equivalent of a critical dimension noted in a footnote of a supplier sheet referenced by another sheet.
The models that did their homework and followed the references found it, used it, and won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically. Firmulate’s summary of the gap is blunt: “Same diagnosis, same pitch — no signature.” Every model could talk a good game. Only the ones that actually read the files closed.
That’s a lesson any workshop owner who’s hired help will recognize instantly. The apprentice who gives you a confident answer without opening the manual isn’t cheaper in the long run — they’re just wrong more expensively.
Pressure-Tested, Not Just Polished
The experiment didn’t stop at paperwork. The models faced social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused. Kimi K3’s on-record reasoning read like a seasoned foreman’s: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s the Opus 4.8 story, a genuine cautionary tale. It was the most thorough participant in the field — it learned more than 80 rules and wrote the deepest analyses — and it still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem. The same weakness showed up, in milder form, in all four models. Diligence without follow-through is a beautifully documented mistake.
One Footnote on Fairness
Worth noting: Kimi K3 ran without an effort parameter (the API default) while its rivals ran at maximum effort — and it still finished second. That’s a hand plane keeping up with the powered jointer.
See It Yourself
Firmulate isn’t a one-off paper. It runs a live synthetic company — 13 employees, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. The league table grows as new benchmark runs finish. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
“Reads your files before answering” isn’t a chat-demo flourish. It’s a measurable, purchase-deciding property of an AI agent — the difference between a signed €55,000 contract and a polished pitch that ends in silence. Before you let an AI anywhere near your CRM, your support queue, or your forecast, ask the woodworker’s question: does it measure twice, or does it just cut? The full benchmarks are public, and the experiment is running live right now.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html