The solution
Long-running tools dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent sits idle. The observation stall is the same tension that drove out-of-order processors — a sequential interface hides work that could be predicted and started early, but a speculative result may only become visible after it and every earlier step have been validated.
TomasuLLM applies the CPU answer to LLM agents. Its runtime executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state.
Across three benchmarks spanning sub-second to minutes-long tool calls, the speedups scaled with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. Across 4,010 audited commit-validation records, the runtime produced zero false accepts — no speculative result was ever committed as if it were real.
Why it worked
Tool latency, not model speed, dominates the wall-clock time of coding agents.
A sequential trajectory interface hides tool work that can be predicted and started before the agent reaches it.
Copy-on-write sandboxes make wrong speculations disposable instead of catastrophic.
Committing in trajectory order after validation keeps the visible behavior sequential while keeping the speedup.
What can be applied
Old hardware tricks travel: when a pipeline stalls on slow steps, predict the work, run it ahead in disposable copies, and gate visibility on validation.
Aftermath
The paper reports results on 100 SWE-bench Verified tasks, 28 Terminal-Bench 2.0 tasks and 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records. Version 2 was posted to arXiv on 1 October 2026.
FOLLOW THE EVIDENCE