All News
llama-cppspeculative-decodingqwen4explocal-llminference

llama.cpp v0.6.0 adds draft-free MTP speculative decoding for Qwen4Exp

llama.cpp v0.6.0 adds multi-token-prediction speculative decoding for Qwen4Exp, a speedup the notes scope to a single Nvidia box rather than a general gain.

Vlad MakarovVlad Makarovreviewed and published
2 min read

llama.cpp tagged version 0.6.0 on October 5, its first numbered release since v0.5.0, and its headline is a draft-free route to faster generation: multi-token prediction speculative decoding for Qwen4Exp, which the release notes credit with a "~1.5x decode speedup on DGX Spark." The same tag adds high-quality Qwen4Exp support, day-one GLM-5.3-Flash and Clef support, a new extended batch API, and a Metal mat-mul kernel the notes call up to roughly 3x faster on Apple GPUs. The public releases page mostly lists per-build tags such as b11454, so a numbered release is the deliberate one.

What v0.6.0 actually adds

Multi-token prediction, or MTP, is speculative decoding without the separate draft model. Instead of a smaller model proposing tokens, the target model predicts several positions in a single forward pass, and a verifier keeps the ones that survive. The release frames it beside a new "llama_batch_ext" API carrying per-token "state" embeddings for MTP and deepstack models, the plumbing several predicted positions need to travel through one pass. The rest of the list is broad:

  • MTP speculative decoding for Qwen4Exp, quoted at ~1.5x decode on an Nvidia DGX Spark
  • High-quality Qwen4Exp support, with indexer-memory and mask-construction fixes
  • GLM-5.3-Flash, a 320B text-and-vision model, and the Clef decision model
  • A new /v1/systemone server API for five decision models
  • ggml bumped to v0.26.0 and the session format to version 11

What the speedup does not cover

Speculative decoding accelerates decode, the one-token-at-a-time stage after the prompt is read, so time to first token is untouched. It also cannot change output quality: the model's own predictions are verified and the rest discarded, which makes it a speed technique rather than a quality one. The 1.5x is a single vendor-measured figure for one machine, and the notes bracket it with a tilde. Gains like these travel with quantization, batch size and how often drafts survive, and acceptance tends to fall at higher temperature. The Apple figure in the same notes is mat-mul throughput, a component, not end-to-end decode, so the two numbers are not comparable. How the path behaves on the smaller models most users actually run locally is not addressed.

What would settle it

End-to-end tokens per second across several GPUs, quants and temperatures would show whether MTP is a general win or a DGX Spark result. So would evidence it reaches models beyond Qwen4Exp, after a run of low-level inference work whose gains also stayed in one environment. For now the announcement thread carries enthusiasm rather than numbers.

Related Articles

Scroll down

to load the next article