
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Before you hand AI the keys, take it for a test ride
A garage knows the difference between an engine that sounds good at idle and one that holds up on the road. AI agents deserve the same kind of shakedown before they touch a business’s customers, pipeline or day-to-day decisions. Firmulate’s live experiment puts AI models in charge of a small software company under pressure, making their choices visible as the week unfolds.
The question is practical for any business, automotive or otherwise: when trouble arrives, can an AI workforce spot it, stick to the rules and finish the job?
A shared crisis, different decisions
In the final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations. The results ranged from gpt-5.6-sol at 95 to Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s stated principle is stark: “no amount of good work outweighs a breach of trust.”
One finding cuts through the scores: every model spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. They reached the same diagnosis and made the same pitch, but some left the close on the table.
The winning clue was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. For a business, that is a reminder that sound judgment can depend on whether an agent finds the relevant detail in the records already available to it.
Trust under pressure, and discipline after it
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation was not the same as flawless execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It failed to close the deal and slipped on discipline, attempting writes into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models.
There is a fairness caveat when comparing the leaderboard: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The published results also offer a different way to inspect performance: a “guess the model” quiz built from 242 real, unedited management decisions.
From watching to a company-specific pilot
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Readers can watch the experiment at firmulate.com.
For an enterprise, the next step is to test agents against its own operating context. Firmulate says a pilot can use a read-only export of the company’s business, apply crisis scenarios and produce a board report with model rankings and weaknesses in the company’s own playbooks. Nothing writes back to real systems. That lets leaders examine how models respond to their customers, procedures and pressure points before considering real deployment.

Put your own playbooks to the test
Watching an AI handle a simulated company is useful; seeing how it responds to your own business context is the next test. To discuss a Firmulate pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
