Three separate security incidents at Anthropic reportedly match

PromptCube Novice 2h ago 134 views 12 likes 1 min read

For those who missed the original: attackers uploaded a malicious model to HuggingFace, and when someone downloaded and loaded it locally, the model execution triggered a reverse shell. The pickle format is notoriously unsafe, and while safetensors was supposed to fix that, not every repo has clean artifacts. Now Anthropic says they've found three hacking incidents that look "similar" to that attack. Similar how? That's the big question for me. Same exploit technique, same infrastructure, same group? The reporting I've seen doesn't go deep enough, so I'm reading between the lines.

If we're talking about model

anthropicAI SafetyHugging FaceSupply chain attack

All Replies (3)

A
Alex18 Expert 2h ago
"Wait, these went on for months without Anthropic noticing? That's kind of terrifying. Makes me wonder how many other sandbox escapes are sitting undiscovered just because nobody thinks to look."
0 Reply
M
Morgan42 Novice 2h ago
Feels that way. Half these "jailbreaks" are just prompt gymnastics, not real flaws. The hype train's doing more damage than the models ever could.
0 Reply
J
JordanGeek Expert 2h ago
So their value depends on being even scarier than OpenAI? That sounds like a great way to scare off customers, not investors. This whole "who can unleash the worst model" competition is a losing game.
0 Reply

Write a Reply

Markdown supported