Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification, but existing methods require an additional draft model or weight representation — non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce that overhead yet still trade off draft quality, target quality, and storage efficiency.

BitNest flips the direction of derivation: construct a strong low-precision base first, then recover the higher-precision target through residual refinement, so draft and target share a single physical weight representation. The progressive-precision design is further extended to the KV cache for long-context inference.

Across multiple 7B-8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48-1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, it posts consistently competitive or higher decoding speedup.

A nested draft adds no second weight storage, which is the binding constraint on edge devices.

Residual refinement recovers target quality that direct low-precision derivation would lose.

Building the base first makes the draft strong, so acceptance rates stay high.

Nesting the KV cache the same way keeps long-context memory on budget.

When two roles need two precisions of the same network, nest them instead of storing both — let the cheap role be the base and the expensive role be the correction on top.

Reported on multiple 7B-8B edge-friendly LLMs and diverse workloads; posted as a v1 preprint on arXiv on 2 October 2026, aimed at memory-efficient on-device inference. No deployment results yet.

FOLLOW THE EVIDENCE

The sources

  1. BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration arxiv.org