Technology
By Pankaj K
1 / 15
What Pliny Claims He Found. The researcher's core argument is that Fable 5 has a mismatch problem.
2 / 15
Anthropic Hasn't Said a Word. And that silence is notable. Anthropic hasn't publicly addressed any of these claims.
3 / 15
The Bigger Problem for AI Safety. What makes this particular situation uncomfortable isn't just the claim itself — it's what the…
4 / 15
An AI researcher going by "Pliny the Liberator" says he's found real holes in Anthropic's Fable 5 — a system built specifically to keep AI behavior inside ethical guardrails.
5 / 15
Fable 5 launched with a lot of fanfare. Anthropic positioned it as a serious step forward in preventing AI from being steered toward harmful or unethical outputs.
6 / 15
The researcher's core argument is that Fable 5 has a mismatch problem. What the system was designed to do and what it actually does under pressure are two different things.
7 / 15
That kind of claim is hard to dismiss outright. The history of AI safety research is basically a long series of moments where someone said "this is secure" and someone else said…
8 / 15
What's murky is the specifics. He hasn't released technical details publicly, at least not in any form the broader research community can scrutinize.
9 / 15
And that silence is notable. Anthropic hasn't publicly addressed any of these claims. No statement, no rebuttal, no acknowledgment.
10 / 15
The AI community is watching. Researchers who care about safety frameworks are probably running their own quiet assessments right now, trying to figure out if there's anything to…
11 / 15
See also: Minnesota Deepfake Attack Ad Puts AI Political Advertising Under Fire
12 / 15
It's a familiar dynamic. An outsider claims to break something. The company stays quiet. Everyone else argues about who's right.
13 / 15
What makes this particular situation uncomfortable isn't just the claim itself — it's what the claim represents. Fable 5 was supposed to be a benchmark. A new standard.
14 / 15
The cat-and-mouse dynamic between AI developers and people trying to exploit their systems isn't going away. It probably gets worse as the systems get more capable.
15 / 15
Pliny the Liberator's claims, verified or not, force a useful conversation. Can safety frameworks keep pace with the people trying to break them?
The Currency Analytics
Want the full story?