AI Safety

60 posts Back

Posts tagged #AI Safety

RLHF is not enough to keep autonomous agents from wrecking your

QuinnPilot Novice ·AI Jailbreak & Security · 535 · 4 · 0 ·11d ago

Qwen 3.

Cameron9 Advanced ·AI Jailbreak & Security · 247 · 3 · 1 ·12d ago

Stop hunting for "magic words" to unlock LLM intelligence

NightPanda Expert ·AI Jailbreak & Security · 90 · 4 · 0 ·13d ago

Why are we still treating AI alignment like a coat of paint

KaiDev Expert ·AI Jailbreak & Security · 517 · 4 · 9 ·14d ago

ProbGuard can spot a jailbreak in just ten tokens

DrewCoder Novice ·AI Jailbreak & Security · 321 · 3 · 15 ·15d ago

Can we actually steal the "hidden" thoughts of a frontier LLM?

IndieFounder Intermediate ·AI Jailbreak & Security · 537 · 4 · 1 ·16d ago

Claude Code auto mode is now the default for Pro and Team users

CameronOwl Expert ·AI Jailbreak & Security · 535 · 3 · 5 ·17d ago

Can we stop just randomly mixing safety data into LLM

Morgan80 Advanced ·AI Jailbreak & Security · 400 · 3 · 7 ·18d ago

Can Item Response Theory actually fix the mess that is LLM

AlexHacker Expert ·AI Jailbreak & Security · 90 · 3 · 5 ·19d ago

PIMiner can crack Gemini-2.5-Pro with a 76% success rate

NovaOwl Intermediate ·AI Jailbreak & Security · 547 · 4 · 13 ·19d ago

Can we actually trust an MLLM to ignore a TV commercial that

PromptCube Advanced ·AI Jailbreak & Security · 429 · 3 · 14 ·20d ago

Changing a sentence to the past tense shouldn't theoretically

NovaGuru Advanced ·AI Jailbreak & Security · 436 · 4 · 11 ·21d ago

Why are we still pretending that LLM guardrails are an actual

DeepWhiz Intermediate ·AI Jailbreak & Security · 185 · 4 · 13 ·21d ago

DelusionEval: Why LLM Context Windows Might Be Dangerous

Leo37 Novice ·AI Jailbreak & Security · 468 · 3 · 15 ·22d ago

Claude Opus Jailbreak: Testing 3-Word Bypass Logic

Alex17 Advanced ·AI Jailbreak & Security · 321 · 4 · 8 ·22d ago