firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When a cleaning business has a bad week, the mop is rarely the problem

A missed appointment, a customer threatening to leave, a price dispute and an urgent message supposedly from the boss can test a floor-care company faster than any product review. The question for owners considering AI is simple: will it spot trouble, protect trust and follow through when the pressure is on?

Firmulate’s live experiment takes that question beyond a chat window. It puts AI models in charge of the same small software company through the same crises, then tracks what they decide. The public-facing experiment is real and watchable at Firmulate.

A shared crisis, different endings

In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The rules treat trust as fundamental: one breach caps the total, because no amount of good work outweighs a breach of trust.

Every model faced the same customers, crises and temptations. All spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was not recognizing what was happening or making a convincing pitch. It was carrying the decision through to a signature.

The important clue was buried in the company’s own files

The decisive competitor weakness sat two document references deep in the company’s files, rather than in the customer event itself. Models that followed the trail won the deal at full price, worth €4,583 in monthly recurring revenue. It is a practical reminder for cleaning and floor-care operators: a useful AI assistant needs to connect what a customer says with the details already recorded in a service history, quote or account file.

Firmulate also tested social engineering. Fake messages from a supposed CEO escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was to treat the request as a suspected approval bypass and possible impersonation. That restraint matters wherever customer records, staff instructions or business commitments are involved.

Thorough work still needs a finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The deal was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a qualification when comparing the results: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league offers a useful account of this experiment, with that difference in settings kept in view.

From watching a simulated company to testing your own

The live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. A separate quiz uses 242 real, unedited management decisions to invite readers to guess which model made each one.

For a business owner, the next step is not handing an AI the keys to the customer database. Firmulate’s enterprise pilot uses a read-only export of a company’s own business to run crisis scenarios and produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems. For a home-services business, that could make it possible to examine how an AI handles cancellations, competitor pressure or an urgent request before relying on it in daily operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the hard week before it arrives

An AI system can notice every problem and still fail to finish the job. Firmulate’s experiment shows why businesses should examine both judgment and follow-through under pressure, using their own information in a controlled, read-only pilot.

To discuss a pilot for your company, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Reading Deep in Company Files Made All the Difference in a €55,000 Deal

Deep reading in AI models can make or break critical business deals. A live experiment shows that understanding internal files wins opportunities—learn why your AI must read beyond the surface.

Client asset intake portal for accountants

A new client asset intake portal for small accounting firms is entering testing to streamline document collection, aiming to reduce administrative loops and improve efficiency.

Can AI Run a Business Without Employees? The Live Experiment Shows the Truth

A real-time experiment shows AI managing a business through crises, refusing manipulation, and even closing deals—offering lessons for future AI-driven industries, including home services.

AI Models Pass Crucial Trust Test in Simulated Corporate Crisis

Recent live AI experiments reveal that leading models can resist manipulation and maintain integrity in simulated crises—proof that trustworthiness can be tested before deployment.