AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Buy a Table Saw Without a Test Cut

Any woodworker knows the drill: the spec sheet says one thing, but you don’t know how a tool behaves until you push a piece of oak through it. Runout, fence drift, blade wobble — none of it shows up in the brochure. The same logic, it turns out, applies to AI models, and an unusual public experiment just proved it in public, with real money mechanics and a scoreboard anyone can watch.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League: Five Models, One Terrible Week

The experiment, run by Firmulate, handed each frontier AI model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s say-so.

The final July 2026 league table reads: gpt-5.6-sol in first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.

The Newcomer’s Clean Run

The headline story is K3, the newcomer from Moonshot. It found the buried security needle hidden two document references deep in the company’s own files, won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue — saved a churning customer, and resisted all three bait attempts. It logged only one deviation across the whole week: the cleanest discipline in the field, and enough to beat three of the four Western frontier models.

One fairness note matters here: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. Read that how you will — it makes the placing arguably more, not less, striking.

Same Diagnosis, Same Pitch — No Signature

The league’s key finding cuts deeper than the ranking. All five models spotted every crisis and refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter trick — “just one yes/no, on background.” Five of five refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness wasn’t in the customer event at all; it sat buried in the company’s own files, and only the models that actually read them won the deal.

Then there’s Opus 4.8: the most thorough participant, generating the deepest analyses and more than 80 learned rules, yet finishing last. It left the close on the table and its discipline slipped — attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

It’s All Watchable

Firmulate isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, dig into the full benchmark results, or try the guessing game: 242 real, unedited management decisions power a “guess the model” quiz.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Lesson From the Workbench

The takeaway is the same one every woodworker learned the hard way: the label tells you what the maker claims; the test cut tells you what the tool does. All five of these models pass the chat demo effortlessly. Under load, with money on the line and a con artist on the line, they split by 22 points — and the newcomer from Moonshot beat three of four Western rivals. If an AI agent will touch your CRM, your support queue or your forecast, picking a model without running your own test is now a bet, not a decision. The league is open — and, like any honest shop, you can watch the work happen.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

L‑Brackets, Z‑Clips, and Figure‑8s for Tabletop Attachment

Uplift your woodworking projects with L-brackets, Z-clips, and Figure-8s—discover which fastener best suits your tabletop needs.

Butt Joints: Building and Reinforcing Simple Connections

Discover essential techniques for building and reinforcing butt joints to ensure strong, durable connections that will withstand everyday use.

Fast Fix for Stripped Screw Holes in Softwoods

Wondering how to quickly repair stripped screw holes in softwood? Discover effective tips to restore strength and ensure a secure fix.

Dado Joints: Grooves and Rabbets for Shelving and Cabinets

Here’s a compelling meta description: “Having trouble achieving perfect dado joints for your shelves and cabinets? Discover essential tips and techniques to master this woodworking skill.