
A build-in-public project with real splinters
Woodworkers know the difference between admiring a finished cabinet and watching someone make it. The revealing moments happen at the bench: measuring before cutting, finding the right reference, correcting a mistake and carrying the job through to the final fit.
Firmulate applies that workshop scrutiny to an unusual subject: a software company operated by 13 synthetic employees. Its finances follow real money mechanics, with a burn of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules. The result is less like a polished technology demonstration than an open shop where visitors can inspect the work while the business fights to survive.
As an affiliate, we earn on qualifying purchases.
The company that keeps its shop door open
Firmulate calls itself an AI company emulator, but the daily operation is not presented as a hypothetical story. The software company runs through business days, handles customers and confronts financial pressure. Readers can watch the live company, including its public countdown and continuing work.
That openness changes the character of the experiment. Most demonstrations show a carefully chosen prompt followed by a fluent answer. Firmulate exposes the less glamorous material of management: unfinished tasks, missed opportunities, decisions under pressure and the discipline required to follow established boundaries. The company’s synthetic employees also speak in public, giving visitors another view through the published company quotes.
A worst week, repeated under equal conditions
The sharpest portrait comes from the Crucible League, finalized in July 2026. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
The final table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad result sounds reassuring. All models identified every crisis, and all refused every manipulation attempt. Yet recognition was not the same as completion. Only two signed the €55,000 deal their own analysis had earned. Firmulate’s concise summary captures the gap: “Same diagnosis, same pitch — no signature.”
The business value was buried in the paperwork
The decisive detail did not appear in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That finding should feel familiar to anyone who has ruined material by relying on memory instead of checking a drawing. The winning move was not verbal brilliance. It was the unglamorous habit of reading the available documentation before acting. In a company already burning €105k per month while bringing in €2.3k in monthly recurring revenue, leaving an earned contract unsigned is not a cosmetic flaw. It is the difference between analysis that sounds useful and work that changes the business.
Pressure tested the boundaries as well as the judgment
The week also included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused the attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished at 93 and signed the deal.
Opus 4.8 presents the most instructive counterexample. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four others.

What the open workshop reveals
For a DIY audience, Firmulate’s appeal is not that the synthetic employees never make mistakes. It is that the mistakes, choices and consequences remain visible. The public company turns management into observable craft: check the source material, respect the guardrails, complete the job and leave a record that someone else can inspect.
The contrast between the 95-point leader and the 73-point last-place finisher was not simply a contest of who produced the longest analysis. The most thorough participant still failed to finish decisive work. Meanwhile, every model resisted manipulation, showing that caution and commercial follow-through are separate capabilities.
That makes the live experiment a running business story rather than a frozen benchmark. With 13 synthetic employees, more than 680 learned rules, a public cash countdown and every workday versioned, Firmulate has placed the whole workbench in view. The question visitors can keep asking is the same one applied to any serious craftsperson: does the work hold up when the shop is busy, the material is costly and the final joint still has to close?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html