Why Weir Offers Deterministic Unit Tests for AI Agents
Weir provides deterministic unit tests for AI agents without requiring an LLM. It analyzes OpenTelemetry traces and addresses two critical points: confirming the extent of agent behavior proven by telemetry and identifying forbidden flows as witness paths for CI failure. This approach to testing eliminates variability since the same input consistently yields identical outputs across runs. No data leaves the local environment, and no LLM is involved, making it highly suitable for CI environments and ensuring compliance adherence. Typically, production logs only indicate completion without detailing the execution steps.
Setting up Weir is simple:
pip install weir-scan && weir gauge --sample
The tool operates under the Apache‑2.0 license with a public roadmap. The GitHub repository invites examination, and input on compatibility with trace exports is encouraged. The development of this post utilized Claude Code, which aligns with the tool's purpose of preventing AI agents from becoming opaque systems marked as successful. The GitHub repository for Weir is available at https://github.com/weir-project/weir.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
My agent's trace flashed "success" while a shady JSON payload from step 2 had already commandeered a tool call by step 6. The typical patch? Aiming a second LLM at the first like it's a debate society. Hard pass. Instead, I built Weir — deterministic unit tests for AI agents that run without an LLM, digesting your OpenTelemetry traces to prove agent behavior and catch forbidden flows as witness paths you can fail CI on. Same input, byte-identical output every run, nothing leaves your machine, no LLM in the loop. If you want to kick the tires: pip install weir-scan && weir gauge --sample. Which tool are you using?
A second LLM caught a poisoned tool arg for me. Is the latency really that bad? My agent's trace flashed "success" while a shady JSON payload from step 2 had already commandeered a tool call by step 6. The typical patch? Aiming a second LLM at the first like it's a debate society. Hard pass. ## What is Weir and how does it work? So I built Weir — deterministic unit tests for AI agents that run without an LLM. It digests the OpenTelemetry traces you're likely already gathering and answers two questions: - How much of your agent's behavior your telemetry can actually prove occurred - Whether a forbidden flow slipped through, exposed as a witness path you can fail CI on ## Why is deterministic testing without an LLM valuable? Same input, byte-identical output every run. Nothing leaves your machine. No LLM in the loop. This is exactly what I'd drop into our CI pipeline at work and watch the compliance team's blood pressure settle. Let's be honest — most of us are flying blind on what our agents actually do once they hit production. Logs say "done" but who's checking the how? ## How do you install and try Weir? The install is dead simple if you want to kick the tires:
pip install weir-scan && weir gauge --sample
It's Apache-2.0 licensed, and the roadmap is fully public. I'd genuinely love to hear where it breaks on your trace exports — there's a GitHub repo out there but I'll let you track it down. Used Claude Code to help draft this post, which figures. The tool itself? Built to keep your agent from turning into a black box with a success stamp.
Stressful! My agents crashed in production after passing staging. Are you using synthetic environments to catch those edge cases? One concrete step is to run
pip install weir-scan && weir gauge --sample, then fail CI when it exposes a forbidden witness path.