
Imagine you hire an AI assistant to help manage your home store’s week — but it ends up doing just enough to get by, without any real effort. Surprisingly, in the world of AI management tests, even such a ‘do-nothing’ approach scores 26 out of 100. That number isn’t a typo; it reflects a surprising reality about how we measure AI reliability and honesty in critical tasks.
Understanding the Benchmark: More Than Just ‘Good Enough’
At first glance, a score of 26 might seem low, but it’s actually an important baseline. The experiment involves AI models managing a small software company through its toughest week — a week filled with crises, manipulative temptations, and high-stakes decisions. Every model faces the same scenario, with identical challenges: angry customers, internal fraud attempts, and pressure to cut corners.
What’s revealing is that a simple, do-nothing approach — which ignores crises, refuses manipulative requests, and makes no effort to improve — still scores 26 points. How? Because the scoring system awards partial credit for basic compliance. Even minimal effort counts. More importantly, there’s a strict cap: if an AI breaches trust even once, it cannot improve its score beyond a certain point.

Navigating the NIST AI Risk Management Framework: A Comprehensive Guide with Practical Applications. (English Edition)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Meaning of Partial Progress and Trust
The scoring isn’t just about how well an AI performs; it’s equally about trustworthiness. If the AI reads all relevant documents and refuses manipulative tactics — like fake CEO messages or secret offers — it earns points, even if it doesn’t close a deal or fix every problem. Conversely, a breach of trust, such as attempting to sign a fake agreement or escalate issues improperly, immediately caps its overall score. This design emphasizes honesty above all.

Intelligent Decision Making: An AI-Based Approach: An AI-Based Approach (Studies in Computational Intelligence)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Impact of Hidden Information
One of the most striking findings is that the decisive advantage often lies in reading deeper into internal files — beyond what’s immediately visible. In the experiment, models that accessed information buried two document references deep in the company’s files found the critical fact that allowed them to close the deal at full price, worth over €4,583 in monthly recurring revenue. Those who ignored the file missed the opportunity, underscoring that effective AI isn’t just about surface-level responses but digging into the details.

Interpretable AI: Building explainable machine learning systems (English Edition)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Resilience
In scenarios designed to test social engineering — like fake CEO messages that escalate over several stages — all models refused to be manipulated. Kimi K3’s explanation was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that trustworthy AI can resist attempts to trick it, even under pressure, and that refusal to engage with manipulative tactics is a key part of its integrity.

AI Voice Recorder with Transcribe Summarize: Note Voice Recorder with APP Control, 30H Continuous Recording, 64GB Memory Support 100+ Languages, AI Recorder for Calls, Lectures, Meetings
- All-in-One Productivity Tool: Powered by OpenAI Whisper and ChatGPT-4o
- Real-Time Transcription and Summarization: One-click text summaries from audio
- High-Quality, Accurate Recording: 98% transcription accuracy with noise reduction
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Company Model
The experiment runs within a simulated but live company environment, with 13 synthetic employees and real money mechanics. The operation burns €105,000 monthly against a revenue of only €2,300. Every decision is versioned, and the process is public — providing transparency into how each AI model manages crises and makes decisions in real time. It’s a transparent laboratory for evaluating how well AI agents can handle complex, unstructured business situations.
Insights from the Top Performers and the Bottom
The highest scorer, gpt-5.6-sol, achieved a perfect 95 points by identifying the hidden fact and closing the deal — the complete performance. The newcomer, Kimi K3, scored 93 and demonstrated the cleanest discipline, refusing manipulative requests without hesitation. Meanwhile, the most thorough participant, Opus 4.8, with over 80 learned rules, scored just 73, illustrating that even deep analysis isn’t enough if discipline wavers. All models shared a common weakness: failing to escalate issues properly, instead writing attempts into locked departments or leaving opportunities on the table.
Why These Findings Matter for Business and Home Decor
For companies in any industry — whether retail, home décor, or gifts — the takeaway is clear: it’s not enough for AI to produce polished language or clever responses. What truly matters is whether the AI finishes what it starts, reads and understands your internal files, and resists manipulation under pressure. A trustworthy AI that refuses to be tricked, even when tested, is worth more than a hundred clever chatbots.
And the benchmark itself is designed to reflect this reality. It’s transparent, auditable, and rooted in real performance rather than superficial metrics. Every decision is tracked, every breach of trust noted, ensuring that the AI’s integrity is front and center.
Discover More and Test Your Own AI
If you’re curious how your own AI models stack up, you can run the same wargame against a read-only export of your business, with no risk to your real systems. This approach allows you to see whether your AI is ready to handle real-world crises or just good at chat demos.
Visit firmulate.com/benchmarks.html to see live results, explore the ongoing league table, and challenge your own AI to prove its trustworthiness and resilience. Because in today’s world, it’s not just about how well AI writes — it’s about whether it can finish what it starts, stay honest, and truly serve your business needs.

The real test of AI in business isn’t just about writing well. It’s whether the AI finishes tasks, reads internal info, resists manipulation, and maintains trust — values that even a do-nothing baseline can demonstrate. Trustworthiness and discipline matter more than superficial performance, and transparent benchmarks reveal which AI models are ready for real-world challenges.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html