Mixture-of-Experts models give fast inference, but training one from scratch or upcycling a dense LLM takes vast compute and corpora — DeepSeek-V3 and Qwen2.5-Max were pretrained on 14.8T and 20T tokens. L0-MoE instead builds a working MoE from a pretrained dense checkpoint and a curated 30B-token corpus.

The mechanism: cluster the corpus into semantic domains with BGE-M3 embeddings and K-means (a cluster confusion matrix orders them by semantic distance), then for each domain train a binary mask over the FFN intermediate dimensions using L0-regularization, annealing the retention ratio from 100% down to target while freezing everything outside the MLPs. That yields 64 experts, each a slice of the original model specialized for one domain.

Dynamic batching then teaches the router in two loops: same-domain batches first so expert selection initializes, then cyclically mixed domains so it learns token-level multi-expert routing.

On Llama-3-8B, Mistral-7B, and Qwen2-7B, L0-MoE reaches 2.0x, 2.1x, and 2.5x inference speedup respectively with benchmark performance comparable to the dense originals — Mistral's variant even averages about 1% better. It outperforms GPTQ quantization, LLM shearing, and RKD+CoT distillation baselines, which speed up but degrade quality; ablations show random, magnitude, OBS, and SVD dimension selection all lose to L0-learned masks.

L0's hard-concrete gates make dimension selection differentiable, so 'which neurons for this domain' is trained rather than searched.

Freezing everything except FFN masks keeps each expert a true slice of the original model, so quality survives.

The cluster confusion matrix orders domains so the router meets familiar ones before mixed ones.

Domain-matched batches bootstrap sequence-level routing before token-level selection is trained on mixed batches.

A dense model already contains its experts; the trick is finding which neurons belong to which domain. Per-domain L0 masks plus a domain curriculum carve them out at a fraction of MoE training cost.

Published in the Proceedings of the 63rd Annual Meeting of the ACL (2025); code and the curated dataset released at github.com/zhangzhenyu13/L0-MOE. The authors position the approach for cost-sensitive industrial deployments where from-scratch MoE training is out of reach.

FOLLOW THE EVIDENCE

The sources

  1. Accelerating Dense LLMs via L0-regularized Mixture-of-Experts arxiv.org