RLHF alone cannot protect against autonomous agents
Training models via RLHF or DPO to ensure "safety" merely instills politeness. This approach suffices for chatbots but lacks structural integrity for LLM agents accessing API endpoints or file systems. When agents mutate databases or execute code, safety must function as a runtime contract enforced by the harness rather than a vague vibe embedded during training.
The safety contract operates on two fronts
Runtime contracts rely on hard boundaries between system execution and model intent, moving beyond simple prompt engineering. The preventive face elevates guardrails to the infrastructure level, employing trajectory monitors, strict permission gates, and sandboxes to intercept disasters before an rm -rf command strikes the disk. The evidential face demands verifiable proof over blind trust in an agent’s claim to have fixed a bug. Committing actions requires hard evidence such as successful test run logs, grounding citations, or file diffs.
Training-time safety loses ground
Research from ICML or NeurIPS reveals an 8x to 12x tilt toward training-time safety compared to deployment-time safety. Yet, documented agent incidents indicate failures stem from agentic loops executing erroneous or hallucinated actions without checks, not from forgotten safety training. Computer security and experimental science demonstrate that systems cannot simply be trained for safety; they need a runtime environment that enforces constraints and demands evidence.
Adopting a Trajectory Schema
The true unit of safety in AI workflows combines trajectory with checkable evidence. The focus must shift from asking if a model is safe to verifying if a trajectory is verifiable. This necessitates a formal Agent Trajectory Schema where the agent’s path consists of compositional gates rather than a raw token stream. Every state change requires an evidence chain, and every step needs monitoring. This shifts the safety burden from the probabilistic nature of the LLM to the deterministic harness. For builders of real-world LLM agents, the harness must serve as the sole source of truth while the model remains an untrusted actor.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Worried. Would a hard-coded verification layer stop the wreckage or just create more bottlenecks? To make it work, task submission should be gated by hard evidence—like file diffs or successful test run logs—so the action isn't committed unless the agent produces the receipt.
I'm definitely feeling the stress. You'll spend all your time debugging those rules instead of actually fixing the model. Training a model with RLHF or DPO to be "safe" is basically just teaching it to be polite. That works fine for a chatbot, but it is structurally insufficient for an LLM agent that has the keys to your file system or API endpoints. If an agent can execute code and mutate databases, safety cannot be a "vibe" instilled during training—it has to be a runtime contract enforced by the harness.
Terrified of agent errors! Read-only API keys help, but they don’t stop every failure—especially if an agent can execute code or mutate databases. Task submission should be gated by hard evidence—file diffs, successful test run logs, or grounding citations.
I’m freaking out after an agent wiped my test DB—what actual safeguards stop this from happening? Just locking down permissions isn’t enough; you need runtime enforcement where the system actively monitors and kills any operation that deviates from a predefined, auditable path before it executes. For example, require the agent to submit a signed diff of changes alongside its request—if it can’t prove the action was safe, the system rejects it outright.