Out-of-order execution for LLM agents, borrowed from CPU design

JamieWolf Advanced 2h ago 480 views 5 likes 2 min read

The latency stall in coding agents — where the model sits idle while a compiler, test suite, or repo command grinds through seconds or minutes of work — mirrors a problem CPU architects solved decades ago. TomasuLLM takes that playbook and applies it to agent trajectories: draft tool calls speculatively, execute them out of order in isolated sandboxes, then commit the results back in sequence only after validating against the committed state.
The paper cites a few hard numbers that make the case concrete:

  • 1.31x mean improvement on 100 SWE-bench Verified tasks
  • 1.35x on 28 Terminal-Bench 2.0 tasks
  • 1.27x matched progress on 18 SWE-Marathon sessions
  • Zero false accepts across 4,010 audited commit-validation records

Those aren't theoretical — they're measured speedups that scale with how long your tools take to run. The longer the tool call, the more headroom there is to speculate ahead.

How the speculation layer works

TomasuLLM doesn't reorder the agent's logic, just its tool execution. It drafts likely next actions based on the current trajectory, runs each in a copy-on-write sandbox so failures don't corrupt state, and traces what each speculative step depends on. Results only get committed in trajectory order after a validation pass checks them against whatever state has been finalized so far. That's the Tomasulo algorithm translated into software: rename registers into sandbox state, track completion tags, and retire in program order.
The key insight is that dependency tracing is what makes out-of-order safe. If step N+2 finishes before step N, it still waits in the buffer until N commits. The sandbox isolation means a failed speculative run never leaks partial state into the main trajectory.

Where it bites back

Speculation only helps when you can predict useful work ahead of time. For tightly coupled reasoning steps where the next tool call depends on the exact output of the previous one, the draft-and-hold pattern adds overhead without parallel payoff. The paper's benchmarks are also all CLI-heavy agent tasks — compilers, tests, repo commands. Whether the same gains hold for API-bound tools or long-context retrieval isn't shown here.
The copy-on-write sandbox model also demands infrastructure you might not have. Each speculative branch needs its own filesystem view and process space, which scales memory and I/O pressure fast if you're drafting aggressively.

What this suggests for the field

This is the kind of cross-pollination that tends to spread. If you're building agent runtimes today, the question isn't whether to speculate but how much sandbox overhead you're willing to carry for the latency win. The zero-false-accept audit record is the part that should make framework authors take notice — correctness validation isn't optional when the world sees results out of order.
One detail worth watching: the speedups are reported as matched-progress gains on SWE-Marathon, not raw task completion. That implies the metric is measuring how far speculatively-executed steps advance the overall goal, not just wall-clock time. Whether that holds up under broader agent workloads is the open question.
Full paper: https://arxiv.org/abs/2609.38201

AI ProgrammingAI Coding

All Replies (0)

Want a live back-and-forth? Join the global AI chat room — login to talk.

No replies yet — be the first!

Write a Reply

Markdown supported