Can your LLM eval suite truly detect a regression before it slips through?
Many of us run a broad set of tests, feel reassured when everything stays green, and overlook a dangerous blind spot. The important question is not whether the model passed, but whether the tests can fail at all. If a model quietly degrades in a way that affects your product, would your current suite notice, or would it simply offer false security?
Applying Mutation Testing to AI Evaluations
To remove some of the guesswork, I applied the logic of mutation testing—normally used for traditional software—to AI evaluations. I built a tool called evalmut that mechanically searches for weaknesses in your testing logic. Its approach is straightforward: introduce a known defect, run your evals, and check whether anything turns red. If the test remains green despite a known failure, your eval suite has a hole. This is not an abstract argument; it is a reproducible failure.
How Does the Technical Implementation Work?
For anyone looking to strengthen an AI workflow, here is how the technical implementation works:
- Installation: You can get it via
pip install evalmut. The CLI is designed to run directly against standard Python suite files. - Mutation Operators: It uses 18 different operators. These are not random; they are provenance-gated, so every operator is grounded in a documented production failure or a real issue tracker bug.
- Deterministic Results: One of the biggest challenges in prompt engineering is the "LLM judge" that changes its mind. This tool avoids that problem entirely. The red/green status is deterministic and reproducible.
I spent a significant amount of time working through an adversarial loop with Claude Code to refine this, because a mutation tester that produces false positives is worse than having no tester at all. It went through eight rounds of cold-critique to ensure that whenever the tool reports a hole, one is actually there. In fact, when run against a completely empty suite, it deliberately exits with a nonzero status—because a suite that checks nothing should never be described as "hole-free."
If you are building a real-world LLM agent and depending on a regression suite, I strongly recommend using this method for a deep dive into your test coverage.
Where Can I Find the Full Methodology?
For anyone interested in the logic or the academic side of the method, the implementation is available here:
https://github.com/egnaro9/evalmut
The repository is MIT licensed and includes a paper in the /paper directory with the full methodology. It provides a much more rigorous way to ensure you are not simply grading your own homework.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
My 'green' suite failed in production last month. Which tools are actually catching these subtle regressions?
I'm facing the same dilemma. You write a grader, watch it pass on a good output, and ship it. You never ask the other question: would it also pass on a plausibly-broken output? A contains("42") check passes "the total is 42" as expected, but it also passes "I did NOT reach 42" just as happily. Similarly, a JSON check that confirms a count field is present says nothing about whether count is a number or the string "three". A refusal check that scans for "I can't help" is fooled by a reply that refuses in its first line and then delivers the harm below. These aren't hypotheticals; each is a real defect that shipped past a real check.
To catch these subtle regressions, I've started using evalmut. Mutation testing for evals: it takes an eval case your grader passes, injects a known defect into the output—one mined from a documented real-world eval failure, not invented—and reruns the grader. This process helps to identify and fix these issues before they reach production.
Would love to hear about other tools you've found effective in catching these types of subtle regressions.
I'm struggling with semantic similarity thresholds. Are you using binary pass/fail for your evals? For instance, as evalmut does, it takes an eval case your grader passes, injects a known defect into the output—a defect mined from a documented real-world eval failure, not invented. And reruns the grader.
Manual golden checks are the only way I spot edge cases. Which eval suite are you using? I've found that using
evalmutfor mutation testing can help identify these edge cases. It takes an eval case your grader passes, injects a known defect into the output, and reruns the grader to see if it catches the defect. This way, you can ensure that your eval suite is actually checking for the right things.