Technology

Story: Researcher “Pliny the Liberator” Claims to Crack Anthropic’s Fable 5…

By Pankaj K

1 / 15

What Pliny Claims He Found. The researcher's core argument is that Fable 5 has a mismatch problem.

2 / 15

Anthropic Hasn't Said a Word. And that silence is notable. Anthropic hasn't publicly addressed any of these claims.

3 / 15

The Bigger Problem for AI Safety. What makes this particular situation uncomfortable isn't just the claim itself — it's what the…

4 / 15

An AI researcher going by "Pliny the Liberator" says he's found real holes in Anthropic's Fable 5 — a system built specifically to keep AI behavior inside ethical guardrails.

5 / 15

Fable 5 launched with a lot of fanfare. Anthropic positioned it as a serious step forward in preventing AI from being steered toward harmful or unethical outputs.

6 / 15

The researcher's core argument is that Fable 5 has a mismatch problem. What the system was designed to do and what it actually does under pressure are two different things.

7 / 15

That kind of claim is hard to dismiss outright. The history of AI safety research is basically a long series of moments where someone said "this is secure" and someone else said…

8 / 15

What's murky is the specifics. He hasn't released technical details publicly, at least not in any form the broader research community can scrutinize.

9 / 15

And that silence is notable. Anthropic hasn't publicly addressed any of these claims. No statement, no rebuttal, no acknowledgment.

10 / 15

The AI community is watching. Researchers who care about safety frameworks are probably running their own quiet assessments right now, trying to figure out if there's anything to…

11 / 15

See also: Minnesota Deepfake Attack Ad Puts AI Political Advertising Under Fire

12 / 15

It's a familiar dynamic. An outsider claims to break something. The company stays quiet. Everyone else argues about who's right.

13 / 15

What makes this particular situation uncomfortable isn't just the claim itself — it's what the claim represents. Fable 5 was supposed to be a benchmark. A new standard.

14 / 15

The cat-and-mouse dynamic between AI developers and people trying to exploit their systems isn't going away. It probably gets worse as the systems get more capable.

15 / 15

Pliny the Liberator's claims, verified or not, force a useful conversation. Can safety frameworks keep pace with the people trying to break them?

The Currency Analytics

Want the full story?