Inkling is a 975B-total/41B-active-parameter mixture-of-experts model. The Cascadia team made it resident on eleven Intel Core Ultra X7 358H AI PCs, each with 64 GB of memory, Arc B390 integrated graphics and gigabit Ethernet — 704 GB of nominal capacity spread across independent address spaces. A custom engine preserves Inkling's routing rules, builds compressed graphs for OpenVINO's fused iGPU primitives, and coordinates FP16 expert computation with FP32 output restoration, fitting six consecutive decoder layers per machine.

Two arrangements made it fast. Dense feed-forward blocks are represented as all-active expert slices, cutting measured dense-layer call time from about 8.1 to 4.5 ms. A streaming pipeline coordinates concurrent generation, reaching 60.29 aggregate decode tokens/s at 88 streams (46.87 over complete serving phases) on a grid tested to 176 streams. Speculative drafting helps single requests: a hybrid proposer uses request-local suffix matches or shared contexts with at least 90 percent agreement, backed by a Qwen3-0.6B CPU draft model — per-request speedups reached 3.28x.

The team also measured the limits honestly. With the context budget raised from the 1,024-position default, real prompts of 1k to 64k tokens recovered the embedded code in all 19 measured answers, and memory holds 512k positions per stream — but first-token time grows as aN+bN-squared because a single-threaded CPU attention loop binds everything (one core busy per machine while the iGPU sat 9-56 percent busy). Head- and key-block parallelism and iGPU attention are named as the next, still unmeasured steps.

MoE sparsity means only 41B of 975B parameters compute per token, so the fleet's real constraint is weight storage, not math.

Keeping weights resident and passing activations between machines avoids streaming hundreds of gigabytes over gigabit Ethernet.

Dense blocks become all-active expert slices so one code path serves both sparse and dense layers on the iGPU.

FP16 expert computation with FP32 output restoration stays close to the deployed numerical path, verified by replaying captured fleet states.

Sparsity changes the economics of scale: a model too big for any one box becomes servable when the weights stay resident and only the activations travel.

The numbers come from the deployed fleet itself: 88 streams proved the best operating point (8 per pipeline group), 176 streams still sustained 57.72 aggregate decode tokens/s, and a start-up probe held 512k positions per stream. Captured-state draft evaluation measured Inkling's shipped multi-token-prediction head at 66.81 percent first-draft agreement, falling only to 64.36 percent with a 65,536-token vocabulary and INT4/INT8 weight grids. The paper frames the setup as an execution and evaluation approach for large sparse models on distributed client systems.

FOLLOW THE EVIDENCE

The sources

  1. Cascadia: Resident 975B MoE Inference on Eleven AI PCs arxiv.org