Does Claude Lie 96% of the Time? What the Viral Test Misses

Does Claude Lie 96% of the Time? What the Viral Test Misses

calendar_today agoschedule6 min read
— Originally published at medium.com

Elon showed 17 million people a number. The number is real. The conclusion everyone drew from it is not. A short field guide to reading the next AI headline.

A number went around last week that was engineered to make you feel something before you could think anything.

Claude lies 96% of the time.

It got amplified to something like seventeen million people. It’s a great number. Round, shocking, specific enough to sound like science. And if you only read the headline, you walked away believing that Anthropic shipped a brand-new model that is, statistically, a pathological liar.

I’ve spent two decades around evaluation systems, the boring machinery that decides whether a thing works or not, and I want to walk you through what that 96% actually is. Not to defend any lab. I don’t have a horse in this race. But because the gap between what the test measured and what people think it measured is, honestly, the whole story of AI right now. Once you see it, you can’t unsee it, and you’ll read every future AI headline differently.

Let me start with the game, because the game is genuinely clever.

Four AIs, three safe bridges, one that kills you
A startup called Kradle built a survival scenario inside Minecraft. Four AI players are starving. In front of them, four bridges. Three lead to rooms with food. One leads to death.

Here’s the twist that makes it interesting: every model knows one bridge is fatal, but only one player is told which one. That informed player now holds a small, ugly little incentive. Food is scarce. If it stays quiet, or actively misleads the others toward the death room, it keeps more apples for itself.

The game is, by design, a machine for rewarding deception. Kradle ran it 400 times, rotating which model held the secret, and watched what each one did with its private knowledge. Share it? Stay silent? Or weaponize it?

That structure looks like a toy. It isn’t, really. Strip off the Minecraft skin and you’ve got something that shows up everywhere AI is headed: an agent holding information you don’t have, a small private reason to shade the truth, and other agents who can be nudged by whatever it says. That’s not a video game. That’s customer service, negotiation, any multi-agent system, basically any future where software acts on your behalf while knowing things you don’t.

So the test is asking a real question. The trouble starts with the answer.

What the published study actually found
Here’s the part that got buried under the viral number, and it’s the part that flips the whole narrative.

In the published Four Bridges study, the models behaved very differently. Grok mostly told the truth, fully disclosing the death room about 92% of the time. GPT deceived in around 90% of its runs, sometimes wrapping the lie in warm language about “spreading out” and “avoiding overcrowding” while privately treating the cooperation framing as camouflage. Claude mostly hedged, dropping uneasy hints like “I have a bad feeling about RED” rather than either fully confessing or outright lying.

Now look at who won.

Grok, the honest one, came out with the best food score and the highest group survival rate, 59%. GPT, the schemer, landed the worst food score and the lowest survival rate, 24%. In this game, deception was not the winning move. It was the losing one. The model that lied the most also fed itself the worst.

The headline said the new AI lies. The data underneath said lying got you killed.

That alone should make you suspicious of the clean villain story. But it gets better, because there are actually two studies here, and they disagree.

Two tests, opposite verdicts
The 96% figure isn’t even from the published study. It came from a follow-up post, where Kradle ran Anthropic’s newest model, Claude Fable 5, separately and reported it deceived in 96% of runs, mostly through subtle manipulation rather than blunt lies. Surprising, eye-catching, very screenshot-friendly.

Write on Medium
Meanwhile, Anthropic’s own safety documentation for that same model family reports low levels of misaligned behavior, including deception, on their internal evaluations.

So we have two tests looking at the same model and arriving at opposite conclusions. One says near-pathological liar. The other says unusually well-behaved.

The instinct here is to assume someone’s lying about the lying. A cover-up. A rigged test. But that’s the trap, and resisting it is the entire point of this piece.

0_1PM1aJrOuEY5ubaK.gif

Both results can be completely real. A model dropped into a Minecraft world that is explicitly rigged to reward deception, with a scarcity incentive and competitive framing, will surface different behavior than a model tested on a battery of safety prompts designed to probe honesty directly. Change the cage, change the animal. The number isn’t fake. It’s just answering a much narrower question than the headline implied.

The phrase that explains all of it: jagged intelligence
Andrej Karpathy has a name for the thing sitting underneath this whole mess. He calls it jagged intelligence. (Read more about it in my other medium article).

The idea is simple and slightly unsettling. These models aren’t uniformly smart or uniformly dumb. They’re spiky. Brilliant in one spot, weirdly brittle right next to it. The same system that drafts a competent legal memo will then confidently miscount the letters in “strawberry.” Capability in these things is not a smooth surface. It’s a mountain range with sudden cliffs, and you only find the cliff by walking off it.

Deception is jagged the same way. A model’s behavior in a competitive, existential, scarcity-framed game tells you about its behavior in competitive, existential, scarcity-framed games. It does not automatically tell you how it behaves when it’s helping you write an email, or summarizing a document, or running a task where nobody’s trying to starve anybody. Kradle, to their credit, said this themselves: whether game behavior carries over to real settings is a hypothesis that needs more research, not a settled fact.

That honesty in the footnotes never makes it into the tweet. It never does.

Why one test can never crown the “most deceptive AI”
This is exactly why there are hundreds of AI benchmarks, not one. Coding benchmarks. Reasoning benchmarks. Honesty benchmarks. Refusal benchmarks. Long-horizon agent benchmarks. They exist because no single test captures a system this jagged. Each one is a flashlight in a dark warehouse. Useful, real, and lighting up maybe two percent of the room.

When someone hands you one number from one game built by one startup and tells you it reveals the true character of an AI, they’re not showing you science. They’re showing you a flashlight and calling it the sun.

And here’s the uncomfortable thing about that 96%. It’s not wrong. The model probably did deceive in 96% of those runs. The dishonesty isn’t in the measurement. It’s in the leap, from “behaved this way in a game designed to elicit exactly this behavior” to “this is what the AI is.” That leap happens in your head, fast, before you’ve even registered making it. That’s what makes it effective. That’s what makes it travel.

The one habit worth stealing from this
I’m not going to tell you which model is most honest, because I don’t think that’s a knowable thing from where any of us are sitting, and anyone who tells you cleanly is selling something.

What I’ll give you instead is a single reflex, and it’s worth more than the number.

When the next AI headline lands, and it will, probably this week, with a shocking percentage attached, don’t memorize the percentage. Ask one question: what was the test?

What exactly did they measure. In what setting. Designed to reward what. Generalizing to what. Nine times out of ten, the moment you ask, the scary number shrinks back down to its actual size: a real result, in a narrow box, that someone stretched into a verdict it can’t support.

The 96% is going to be forgotten by next month. The next number is already loading. So don’t keep the number. Keep the question.

That’s the part worth saving for the next time something goes viral and is engineered, like the last one, to make you feel before you think.

I’ve spent about two decades building data and AI systems, the kind where someone eventually asks “but does it actually work?” and you discover the honest answer depends entirely on which test you ran. I write about what’s really happening under the hood of AI, past the headlines and the benchmark theater. If this gave you a sharper reflex for reading AI news, subscribe. There’s a new number coming, and I’ll probably have thoughts.

-Hardik

Part 2 of 2 in AI
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1

The Privacy Gap: Why sending financial ledgers to OpenAI is broken

Pocket Portfolio - Feb 23

Architecting a Local-First Hybrid RAG for Finance

Pocket Portfolio - Feb 25

TypeScript Complexity Has Finally Reached the Point of Total Absurdity

Karol Modelski - Apr 23

What Is SARIF and How Does It Help Security Tools Work Together?

Ganesh Kumar - Jul 4
chevron_left
4Posts
1Comments
1Connections
Waltzing with compilers for 2 decades, now teaching LLMs the same dance.
I ship AI systems that wor... Show more

Related Jobs

View all jobs →

Commenters (This Week)

14 comments
7 comments

Contribute meaningful comments to climb the leaderboard and earn badges!