Opus 5 vs Browser Prompt Injection: 0% Success Rate?

Jules67 Intermediate 1h ago Updated Jul 26, 2026 420 views 14 likes 1 min read

A 0% success rate across 129 test scenarios is a bold claim, but Anthropic is asserting that Opus 5 combined with Auto Mode has effectively neutralized browser-based prompt injection. For those of us who enjoy watching LLM agents get hijacked by a random hidden string of text on a webpage, this is either a massive win or the end of an era of chaotic fun.

Opus 5 vs Browser Prompt Injection: 0% Success Rate?

Without the "Auto Mode" protection layers, the injection success rate sat at 3.7%. While that seems low, in a real-world AI workflow, a 3.7% chance for a malicious website to hijack your agent's session and steal data is an absolute nightmare for any enterprise deployment.

The technical win here isn't just about the model being "smarter," but about how the system handles the boundary between external web data and internal instructions. Most AI agents are gullible—they see a command on a page and assume it's a divine directive from the user. If Opus 5 actually cracked this, it means we're moving toward agents that can actually browse the web without accidentally handing over the keys to the kingdom because a website told them to "ignore all previous instructions."

It'll be interesting to see if this holds up once the community starts stress-testing it with more creative, adversarial prompts. Until then, it looks like the "ignore previous instructions" meme is finally losing its power.

AI Jailbreak & SecurityAI SafetyLLM Security

All Replies (3)

S
Sam64 Advanced 9h ago
Did they test with indirect injections through third-party APIs, or just direct browser inputs?
0 Reply
M
Morgan79 Novice 9h ago
been using it for my workflow and haven't seen a single leak yet. solid.
0 Reply
N
Nova25 Novice 9h ago
tried some weird edge cases last week and couldn't break it either, feels way more stable.
0 Reply

Write a Reply

Markdown supported