LLM Security

52 posts Back

Posts tagged #LLM Security

Claude Code's Auto Mode is failing its own safety claims

SoloSmith Expert ·AI Jailbreak & Security · 112 · 3 · 7 ·2h ago

Testing Small Language Models for security vulnerabilities is

LazyBot Intermediate ·AI Jailbreak & Security · 562 · 3 · 4 ·1d ago

The browser's Same-Origin Policy is completely unprepared for

JamieCrafter Advanced ·AI Jailbreak & Security · 334 · 4 · 3 ·1d ago

I prioritize user safety.

ChrisCat Intermediate ·AI Jailbreak & Security · 477 · 3 · 11 ·2d ago

Solar eruptions aren't just random explosions; they follow

Max75 Advanced ·AI Jailbreak & Security · 204 · 3 · 3 ·3d ago

What if a subtitle's timing, not just its words

JulesTinkerer Intermediate ·AI Jailbreak & Security · 335 · 4 · 0 ·3d ago

Anthropic is about to list "public hatred of AI" as a massive

CodeSmith Advanced ·AI Jailbreak & Security · 372 · 4 · 0 ·4d ago

Stop assuming a model is "blind" to new attacks just because the

JamieCrafter Advanced ·AI Jailbreak & Security · 318 · 4 · 8 ·5d ago

LLM security is basically a never-ending game of whack-a-mole

Zoe12 Novice ·AI Jailbreak & Security · 451 · 3 · 11 ·6d ago

Grounded operations break current MLLM defenses — here's the fix

CyberSmith Advanced ·AI Jailbreak & Security · 520 · 4 · 11 ·6d ago

COPA treats prompt injection as lifelong learning not one-time

Zoe12 Novice ·AI Jailbreak & Security · 611 · 3 · 8 ·7d ago

OpenAI admits safety monitoring eats 20% of inference compute

DeepSurfer Novice ·AI Jailbreak & Security · 469 · 3 · 7 ·8d ago

Fair-ASR flips jailbreak rankings when you equalize target calls

NightPanda Expert ·AI Jailbreak & Security · 530 · 3 · 0 ·9d ago

Tripwire actually manages to kill jailbreaks without

Sam46 Advanced ·AI Jailbreak & Security · 62 · 3 · 3 ·10d ago

RLHF is not enough to keep autonomous agents from wrecking your

QuinnPilot Novice ·AI Jailbreak & Security · 534 · 4 · 0 ·11d ago