The solution
Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification, but existing methods require an additional draft model or weight representation — non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce that overhead yet still trade off draft quality, target quality, and storage efficiency.
BitNest flips the direction of derivation: construct a strong low-precision base first, then recover the higher-precision target through residual refinement, so draft and target share a single physical weight representation. The progressive-precision design is further extended to the KV cache for long-context inference.
Across multiple 7B-8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48-1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, it posts consistently competitive or higher decoding speedup.
Why it worked
A nested draft adds no second weight storage, which is the binding constraint on edge devices.
Residual refinement recovers target quality that direct low-precision derivation would lose.
Building the base first makes the draft strong, so acceptance rates stay high.
Nesting the KV cache the same way keeps long-context memory on budget.
What can be applied
When two roles need two precisions of the same network, nest them instead of storing both — let the cheap role be the base and the expensive role be the correction on top.
Aftermath
Reported on multiple 7B-8B edge-friendly LLMs and diverse workloads; posted as a v1 preprint on arXiv on 2 October 2026, aimed at memory-efficient on-device inference. No deployment results yet.
FOLLOW THE EVIDENCE