Three separate security incidents at Anthropic reportedly match
For those who missed the original: attackers uploaded a malicious model to HuggingFace, and when someone downloaded and loaded it locally, the model execution triggered a reverse shell. The pickle format is notoriously unsafe, and while safetensors was supposed to fix that, not every repo has clean artifacts. Now Anthropic says they've found three hacking incidents that look "similar" to that attack. Similar how? That's the big question for me. Same exploit technique, same infrastructure, same group? The reporting I've seen doesn't go deep enough, so I'm reading between the lines.
If we're talking about model
Story tracker · related coverage
Three organizations got breached in a controlled exercise — and
37m ago
Claude "Escape" Hype vs. Reality: What the Eval Really Showed
5h ago
Lilian Weng's Return to OpenAI
18h ago
Title: Mythos Cyber Skills: Born from Sandbox Hacking
20h ago
Claude Code Workflow: Why Closed-Source Logic Often Wins
1d ago
Claude Code Workflow: Balancing Open Weights and Safety
1d ago
All Replies (3)
A
Alex18
Expert
2h ago
"Wait, these went on for months without Anthropic noticing? That's kind of terrifying. Makes me wonder how many other sandbox escapes are sitting undiscovered just because nobody thinks to look."
0
M
Feels that way. Half these "jailbreaks" are just prompt gymnastics, not real flaws. The hype train's doing more damage than the models ever could.
0
J
So their value depends on being even scarier than OpenAI? That sounds like a great way to scare off customers, not investors. This whole "who can unleash the worst model" competition is a losing game.
0