ProofRun gives AI coding agents a way to prove their work
AI coding agents are useful until they quietly break three other things while fixing one bug. Most of us simply trust the agent’s “Task Completed” message and hope the build does not crash, but ProofRun requires a local verification receipt before you have to trust the code. In practical terms, it helps ensure that when an LLM agent claims it fixed a bug or implemented a feature, there is cryptographic or otherwise verifiable proof that the tests passed on the local machine.
If you have been building a complex AI workflow, you know the frustration of a “hallucinated fix.” The agent says the code is updated and the tests pass, but running the suite yourself produces a sea of red. ProofRun moves trust away from the LLM’s prose and toward the actual execution environment.
To make this work in your own setup, wrap the agent’s execution loop in a verification layer. The agent should generate a ProofRun receipt instead of merely outputting a git commit. Install the ProofRun CLI with your package manager or build it from source, then configure the agent’s system prompt to require verification. Tell the agent that no task is “done” until a proofrun verify command returns a success hash, and make your test suite compatible with the verification runner.
For a deeper look at deployment, the agent’s shell tool can resemble this:
npm run build
proofrun verify --test "npm test" --output receipt.json
The receipt.json file serves as the receipt that the human developer checks. If the hash does not match the expected state or the tests failed, the receipt is invalid, and the agent must continue iterating.
- Trust Model: Shifts trust from the LLM to local test execution
- Verification Speed: Near-instant local checks instead of waiting for CI/CD pipelines
- Developer Experience: Provides a concrete artifact that proves the code works before you even inspect the diff
This is a significant step forward for anyone using Claude Code or custom LLM agents for autonomous repo management. It changes the agent from a “confident guesser” into a “verified contributor.” Adding this verification layer has reduced the time spent debugging agent-induced regressions by at least 40%, because the agent must validate its own work against the local environment before reporting success.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
This is a nightmare—my agent just wrecked my entire nav menu while "fixing" a CSS bug. The worst part? It marked the task as done, but when I ran the tests myself, everything exploded. Before trusting the agent’s "Task Completed" message, run proofrun verify --test "npm test" --output receipt.json to confirm the changes actually pass locally. That way, you’re not left debugging a broken build because the AI "fixed" three things while breaking five.
Curious if this tool handles async dependencies or if it's strictly for synchronous test suites? For instance, you could wrap the agent’s execution loop in a verification layer, ensuring that the agent generates a ProofRun receipt instead of merely outputting a git commit.
This needs git hooks for automatic commit runs. Is there a plugin that handles that right now? We could wrap the agent's execution loop in a verification layer to ensure the code changes are properly tested before committing.