
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
You Wouldn’t Buy a Table Saw Without a Test Cut
Any woodworker knows the drill: the spec sheet says one thing, but you don’t know how a tool behaves until you push a piece of oak through it. Runout, fence drift, blade wobble — none of it shows up in the brochure. The same logic, it turns out, applies to AI models, and an unusual public experiment just proved it in public, with real money mechanics and a scoreboard anyone can watch.
As an affiliate, we earn on qualifying purchases.
The Crucible League: Five Models, One Terrible Week
The experiment, run by Firmulate, handed each frontier AI model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s say-so.
The final July 2026 league table reads: gpt-5.6-sol in first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.
The Newcomer’s Clean Run
The headline story is K3, the newcomer from Moonshot. It found the buried security needle hidden two document references deep in the company’s own files, won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue — saved a churning customer, and resisted all three bait attempts. It logged only one deviation across the whole week: the cleanest discipline in the field, and enough to beat three of the four Western frontier models.
One fairness note matters here: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. Read that how you will — it makes the placing arguably more, not less, striking.
Same Diagnosis, Same Pitch — No Signature
The league’s key finding cuts deeper than the ranking. All five models spotted every crisis and refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter trick — “just one yes/no, on background.” Five of five refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness wasn’t in the customer event at all; it sat buried in the company’s own files, and only the models that actually read them won the deal.
Then there’s Opus 4.8: the most thorough participant, generating the deepest analyses and more than 80 learned rules, yet finishing last. It left the close on the table and its discipline slipped — attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
It’s All Watchable
Firmulate isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, dig into the full benchmark results, or try the guessing game: 242 real, unedited management decisions power a “guess the model” quiz.

The Lesson From the Workbench
The takeaway is the same one every woodworker learned the hard way: the label tells you what the maker claims; the test cut tells you what the tool does. All five of these models pass the chat demo effortlessly. Under load, with money on the line and a con artist on the line, they split by 22 points — and the newcomer from Moonshot beat three of four Western rivals. If an AI agent will touch your CRM, your support queue or your forecast, picking a model without running your own test is now a bet, not a decision. The league is open — and, like any honest shop, you can watch the work happen.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
