All News
llama-cppmoelocal-llminferencegpu-cache

llama.cpp merges a GPU cache for MoE experts parked in host memory

llama.cpp merged a GPU cache for MoE experts kept in host memory. The speedups shared for 10GB cards are user reports, not benchmarks, and only a first step.

Vlad MakarovVlad Makarovreviewed and published
5 min read
llama.cpp merges a GPU cache for MoE experts parked in host memory

llama.cpp now keeps some Mixture-of-Experts weights on the GPU even when the model does not fit in VRAM. The project merged pull request #29887 on 7 October, as commit d6cf9ac, adding a GPU-side cache for MoE experts whose weights live in host memory and gated by two new flags, -cmoe and --moe-cache-mib. For anyone running a large MoE model on a 10 or 16 GB card, it is the most concrete change to local MoE inference in months. It is also a first step, and nearly every number circulating for it is a user report rather than a measured benchmark.

The mechanism, in plain language

MoE models carry a large total parameter count but activate only a fraction of their experts per token. The usual way to fit one on a small card is to leave the inactive experts in system RAM and pull them to the GPU on demand, but that copy is an I/O bottleneck: every token has to stream the routed experts' weights over PCIe, and decode slows to a crawl. PR #29887, a port of work from Tether's qvac-fabric project, instead keeps a cache of the most frequently used experts in VRAM, so the expert matmuls — the MUL_MAT_ID operation — run on the GPU and only the cache misses get uploaded. The cache only serves small batches of at most 32 tokens; larger batches fall back to the existing path, and layers with different expert layouts get separate banks. The author's note says roughly 10% of total expert size proved a good default, and the feature is switched on with --moe-cache-mib. It is opt-in and changes nothing when the flag is absent. The premise is temporal locality: even when experts are chosen roughly evenly across a long session, the experts picked from one token to the next cluster enough that a modest cache captures most of the traffic.

What the reported numbers actually say

Most of the figures attached to the patch come from the pull request itself or from users in the r/LocalLLaMA thread. The author's own table on Qwen3.8-Flash-Next Q4 puts decode on an RTX 4090 at 25.0 tokens per second stock and 40.7 with -cmoe, and on an RTX 5090 at 30.8 rising to 67.8 — a 2.2x jump at an 89% cache hit rate. Those are single-machine, single-prompt measurements, and the hardware is far above the cards the feature is meant to help. The consumer-grade numbers below all come from the thread, and each is one user on one machine.

GPUModel (quant)Reported decode
RTX 3080 10GBQwen3.6-35B-A3B Q435 to 47 tok/s
RTX 3060 12GBQwen3.6-35B-A3B Q435 to 45 tok/s
RTX PRO 4500 32GBQwen3.8-Flash-Next Q822 to 35 tok/s

The RTX 3080 figures even moved between tests by the same person, from 34 to 53 tokens per second in one run and 35 to 47 with a 4 GB cache in another, with prefill slipping from about 539 to 453 tokens per second. One widely repeated summary of the thread claims 33 to 56 tokens per second with 8 GB of VRAM; this article could not confirm that figure against the comments, so it stands only as an unverified community claim.

The reports cut both ways

The gains are not uniform, and the thread is candid about it. A user on an Intel B580 saw decode fall from 28 to 14.5 tokens per second with a 600 MiB cache; another on an 8 GB laptop found the cache far slower at 2 to 4 GB. One commenter noted the cache only pays off above roughly a 90% hit rate, and that a pool small enough to hit 50% can lose to no cache at all. Systems with slow PCIe or slow RAM, close to the configurations the feature is meant to rescue, are also where the copy cost bites hardest. Commenters also reported that the sweet spot for cache size varies by model — around 7 GB for one Gemma variant, a couple of GB for a Qwen — so no single recommended default fits everyone. That is expected for a community patch on mixed hardware, but it makes the headline that MoE is now fast on cheap cards a conditional, not a rule.

The follow-up, and the direction of travel

A day after the merge, a contributor opened issue #29949 proposing a different design: a GPU-resident LRU cache with a fixed-size per-layer pool of slots, exposed as --moe-lru-slots N, where the eviction decision and the copy both run on the GPU inside the compute graph, and the expert weights sit in pinned host memory. The author says his implementation measured about 1.5x faster than #29887 in same-machine, same-prompt tests — 31.2 against 20.1 tokens per second at an equal cache budget — and reaches 37 to 50 tokens per second decode on a 24 GB Radeon 7900 XTX, up from 9.7 with no cache. The same write-up points to an earlier pull request, #27861, that proposed a GPU-resident LRU in August and has seen no author updates for weeks.

What would settle it

The honest state of play is that llama.cpp now has one merged expert cache whose speedups are user-reported on mixed hardware, and one unmerged follow-up claiming to beat it. What would move the claims from anecdote to fact is a standard benchmark — fixed models, quants and prompts, published across several GPUs, with cache hit rates and prefill costs reported next to decode. Also missing is multi-GPU support, a limitation several commenters flagged and which the follow-up only partly addresses, and any sign that other engines are adopting the technique. Independent runs on the same hardware, by people with no stake in either patch, are the missing piece.

For now the practical advice a 10 GB owner will keep hearing is the one the thread repeats: turn on -cmoe, set --moe-cache-mib generously, and confirm the cache is actually hitting before trusting the gain. The direction is right, and it lands in a busy week for the project that follows the v0.6.0 MTP release and the same audience chasing cheap VRAM rigs. The evidence is still other people's.

Related Articles

Scroll down

to load the next article