AI Safety

26 posts Back

Posts tagged #AI Safety

Kimi K3 Abliterated: A Tool for Blackbox Red Teaming

JordanSurfer Intermediate ·AI Jailbreak & Security · 228 · 3 · 6 ·1d ago

Moderation APIs vs LLM Judges: The Policy Gap

PromptCube Novice ·AI Jailbreak & Security · 290 · 3 · 6 ·1d ago

LLM Safety: Why ASR is a Bad Metric for Defenders

Pat31 Advanced ·AI Jailbreak & Security · 179 · 3 · 9 ·1d ago

Preemptive Hardening for Agentic LLM Security

Max75 Advanced ·AI Jailbreak & Security · 291 · 4 · 6 ·1d ago

DARWIN: Evolving LLM Jailbreak Framework

Dev26 Expert ·AI Jailbreak & Security · 406 · 3 · 4 ·1d ago

OpenAI vs Hugging Face: The Accidental Breach

ZenMaster Expert ·AI Jailbreak & Security · 610 · 3 · 2 ·1d ago

Prismata: Stopping Cross-Site Prompt Injection

Jamie89 Intermediate ·AI Jailbreak & Security · 226 · 3 · 14 ·2d ago

DNS Exfiltration via macOS Terminal ANSI Codes

Jules45 Expert ·AI Jailbreak & Security · 291 · 4 · 14 ·2d ago

OpenClaw Defense: Handling Indirect Prompt Injection

Jamie67 Novice ·AI Jailbreak & Security · 200 · 4 · 13 ·2d ago

AI Safety Leadership Shake-up: Commerce Dept. Exit

TurboFox Novice ·AI Jailbreak & Security · 373 · 3 · 10 ·2d ago

ReasonGate: Stopping Prompt Injection with Explainability

Jordan37 Intermediate ·AI Jailbreak & Security · 509 · 3 · 8 ·2d ago