AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine running a busy home decor online store facing a week of relentless crises: supplier delays, customer complaints, and security threats. Now, what if your AI assistant had to navigate this chaos, make tough calls, and close sales—just like a seasoned manager? Welcome to the world of live AI management testing, where the question isn’t just about how well these models chat, but whether they can truly lead a company through its toughest moments.

The Experiment: Putting AI to the Test in Real Business Crisis

Firmulate, a pioneering company in AI management simulation, recently hosted a groundbreaking live experiment. Four of the most advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—were tasked with running a small software firm through its worst week. Every detail was replicated from real life: identical customers, simultaneous crises, and the temptations to cut corners or manipulate the system.

Each model’s decisions were fully versioned and auditable, meaning every choice could be traced back and analyzed. This setup allowed observers to evaluate if these AI managers could identify crises, handle them ethically, and ultimately secure a crucial €55,000 deal.

Business Intelligence in the Age of AI: Modern Data Warehousing, Analytics and AI-driven Decision-Making

Business Intelligence in the Age of AI: Modern Data Warehousing, Analytics and AI-driven Decision-Making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Models Discovered and How They Reacted

All models showed remarkable competence—they identified every crisis and refused every attempt at manipulation or undue influence. This included social engineering tactics such as staged CEO messages and a reporter inquiry that would normally tempt quick, unethical responses.

However, when it came to closing the big deal, only two models succeeded in signing the contract based on their own analyses. The other two, despite similar diagnoses and pitches, left the opportunity on the table. Interestingly, the decisive advantage was in reading company files deeply hidden within internal documents. The models that identified this buried fact won the deal at full price, adding €4,583 MRR to the company’s revenue.

AI Native Product Design and Intelligent Automation Business Models: Strategic Insights and Frameworks for Building Intelligent, Scalable, and Ethical AI-Driven Products and Business Models

AI Native Product Design and Intelligent Automation Business Models: Strategic Insights and Frameworks for Building Intelligent, Scalable, and Ethical AI-Driven Products and Business Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Personality Profiles: How Different AI Styles Impact Decision-Making

The models demonstrated distinct ‘management personalities’—a fascinating insight into how AI can be tuned. Opus 4.8, the most thorough and analytical, performed best in identifying hidden opportunities but failed to close the deal because it left the final step unexecuted. Its discipline slipped when faced with the temptation to defer or escalate decisions improperly. Meanwhile, Kimi K3, running at a default (lower) effort parameter, showed the cleanest discipline and closed the deal successfully but without the deepest insights.

The other models, featuring different balances of thoroughness and decisiveness, exhibited varying degrees of slip-ups and process slips, highlighting how personality settings influence management style and effectiveness.

AI-Driven Customer Service Revolution: Enhance User Experience & Cut Costs with ChatGPT: Discover Case Studies of Successful AI Implementations in Customer Support for Business Growth

AI-Driven Customer Service Revolution: Enhance User Experience & Cut Costs with ChatGPT: Discover Case Studies of Successful AI Implementations in Customer Support for Business Growth

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust, Ethics, and Resilience Under Pressure

Beyond the deal-closing metrics, all four models showcased an unwavering stance against social engineering attempts, refusing to be duped into unethical shortcuts. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation,” exemplifying their built-in ethical safeguards.

Readiris Business PDF Management Software - PME und Unternehmen

Readiris Business PDF Management Software – PME und Unternehmen

  • Comprehensive PDF Management: Create, edit, comment, share, and sign PDFs
  • Easy Document Conversion: Convert images and PDFs to Word and other formats
  • Multilingual OCR Support: Recognizes 138 languages for importing and scanning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business: A Real Software Company Losing Money

This experiment isn’t just theoretical. The live company run by these AI models is real—comprising 13 synthetic employees managing actual money mechanics. It burns €105,000 monthly against a modest €2,300 MRR, with a public cash countdown and over 680 self-learned rules guiding daily decisions. Every day, decisions are versioned, and the entire process is transparent and watchable at firmulate.com/live.

The Surprising Takeaway: More Than Just Chat Skills

In the world of AI management, success isn’t about how smoothly the models can generate conversations or reports. It’s about whether they can read hidden information, stay honest under pressure, and see opportunities others miss. The experiment revealed that the strongest models—like gpt-5.6-sol—could uncover buried facts and close deals at full price, illustrating a critical difference in management quality that isn’t visible in typical AI demos.

Implications for Business Leaders and AI Developers

This live test underscores a pivotal point: deploying AI in decision-making roles requires understanding their personalities and ethical boundaries. Choosing the right model isn’t just about raw scores but about aligning the AI’s management style with your company’s values and risk appetite.

Want to see how your AI tools stack up? You can run your own business wargame against a read-only export of your operations—nothing ever touches your real systems. Explore the possibilities at firmulate.com/pilot.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Danke mit Humor: Lustige Dankessprüche, die Mama zum Lächeln bringen

Entdecken Sie witzige Dankessprüche, die Ihre Mutter zum Lächeln bringen und sie wirklich schätzen lassen – denn ein bisschen Humor kann jede Geste der Dankbarkeit aufhellen.

Die besten Dankessprüche für Mütter, die nie im Rampenlicht stehen

Absolut wertschätze ich die herzlichen Dankessprüche für Mütter, deren stille Stärke unendliche Anerkennung verdient.

Why Even a Do-Nothing AI Gets a 26 in Business Benchmarks

AI benchmarks reveal that even a do-nothing approach scores 26 points, emphasizing the importance of trust, thoroughness, and integrity in real-world business AI performance.

Danke-Zitate für Mütter, die nicht viel sagen

Ich habe aufrichtig Wege entdeckt, Mütter zu schätzen, die leise sprechen, und du wirst sehen, wie diese Zitate ihre stille Liebe wirklich ehren können.