Why "uncertainty stacking" breaks self-adaptive systems built on LLMs and agents
A new position paper on arXiv (2610.08881) argues that the self-adaptive systems community has been treating uncertainty as a single-layer problem when it actually compounds across the stack. The authors—eight researchers from institutions including IT University of Copenhagen, Universidad de los Andes, and University of Toronto—lay out why GenAI components like LLMs and agentic subsystems don't just add new failure modes, they amplify existing ones through interaction effects.
The core claim: uncertainty from probabilistic outputs, hallucinations, context window limits, and stale memory doesn't stay isolated. It propagates. An LLM serving layer might hallucinate a config value, which then feeds into an agentic decision loop, which then triggers a bad adaptation at the system level. The paper's framework attempts to map these interactions rather than treat each source in isolation.
What the mitigation function actually does
The proposed mitigation function is the centerpiece. It takes five inputs: uncertainty interaction categories, the involved uncertainty types, system design characteristics, criticality, and mitigating architectural tactics. The output maps to impacted quality attributes and the tradeoffs you inherit by choosing one tactic over another.
That's the part worth paying attention to. Most existing approaches ask "how do I reduce hallucination?" or "how do I handle a stale context?" This paper asks "how does reducing one uncertainty source change the risk profile of another?" That's a different optimization problem entirely.
The lifecycle angle is the hard constraint
The authors argue you can't decouple mitigation from runtime monitoring and analysis. That's a direct challenge to design-time-only approaches. If you only handle uncertainty during requirements or architecture phases, you miss what happens when the system actually runs and the LLM starts producing outputs you didn't anticipate.
This implies your mitigation strategy needs to be a feedback loop, not a one-shot decision. You need to observe how uncertainty manifests at runtime, feed that back into the mitigation function, and re-evaluate your tactics continuously.
Three use cases show where this bites
The paper applies the framework to three scenarios: LLM serving, AI-driven software operations, and agentic cloud management. These aren't toy examples—they're the three areas where production GenAI systems actually fail today.
For LLM serving, the uncertainty interaction is between context window limits and probabilistic outputs. You can't fix one without affecting the other. For agentic cloud management, the issue is compounding: each agent's uncertainty feeds into the next agent's decision, so error rates multiply rather than add.
What to do with this
If you're building self-adaptive systems with GenAI components, the practical takeaway is to catalog your uncertainty sources explicitly and map how they interact before writing any mitigation code. Don't assume a hallucination guardrail in one layer protects downstream layers. The paper's framework gives you a vocabulary for that mapping, but it's still preliminary—expect to adapt it to your specific system.
The full framework and the three case studies are in the preprint at https://arxiv.org/abs/2610.08881. Worth a read if you're in the self-adaptive systems space, less relevant if you're just calling an API and handling failures manually.
The paper's point about compounding uncertainty is spot on—I've seen this firsthand with a multi-agent system I built for a client. The more agents I added, the harder it was to track where errors were coming from.