All News
llama-cppinferencelocal-llmoptimizationn-gram

llama.cpp prompt lookup drafting gets 42x faster, in a fork

Hayder Tirmazi's llama.cpp patches speed up n-gram prompt lookup drafting up to 42x, and 140x with Daniel Lemire's follow-up. A costlier stage sits behind it.

Vlad MakarovVlad Makarovreviewed and published
2 min read
llama.cpp prompt lookup drafting gets 42x faster, in a fork

Hayder Tirmazi has published a set of performance patches that make prompt lookup drafting in llama.cpp up to 42x faster. A follow-up pull request from Daniel Lemire pushes the same loop to a claimed 140x, with peak memory down as much as 2.6x. Prompt lookup is a special case of speculative decoding whose draft model is an n-gram model, and llama.cpp, vLLM and Hugging Face's transformers all support it. All five patches are open pull requests in Tirmazi's fork of the project, and none has landed upstream.

What the patches change

llama.cpp keeps three n-gram caches, context, dynamic and static, as nested hash maps. Tirmazi reads the inner maps by reference instead of copying them on every drafting step, then swaps the outer std::unordered_map for ankerl::unordered_dense, replaces the inner maps with sorted vectors behind a fixed-length binary search, and backs the static cache with Lemire's constmap. The static cache is the piece that degrades worst as the corpus grows. Lemire's separate patch skips scoring candidates that cannot clear llama.cpp's own a/p thresholds. His measurements, on an M4 Pro:

  • Drafting latency per drafted token, WikiText-103, median of three runs
  • Upstream: 8.54 microseconds with no corpus, rising to 165.48 at 541 MB
  • All four changes: 0.89 to 3.98 microseconds
  • With Lemire's pre-check: 0.45 to 1.18 microseconds
  • Static cache load: 5.49s to 0.23s; peak memory 3.47 GB to 1.31 GB

What 42x does not measure

Drafting is one stage of generation, and it sits in front of a model forward pass this benchmark never runs. Prompt lookup proposes tokens; the model then verifies them, so the end-to-end gain depends on how many drafts are accepted and on what verification costs. Tirmazi does not report tokens per second, time to first token, or any real workload. His tool replays WikiText-103 test text as if it were model output, which measures the lookup loop rather than generation, and the headline numbers come from the 541 MB corpus, where the upstream loop is slowest. Acceptance rates are unchanged, because he makes no algorithmic change to the drafting rule, so this is the cheap half of the pipeline, optimized. The figures are also self-reported from one Apple M4 Pro with 14 cores and 48 GB, and the fastest result is not in the llama.cpp that people install.

What would settle it

End-to-end tokens per second and time to first token, before and after, on a real model and a repetitive real workload such as code editing or retrieval-heavy prompts, would settle it. So would a merge upstream into the project past 100,000 stars. For now, the Hacker News thread drew 58 points and 9 comments, while the Reddit thread was larger, at roughly 410 points.

Related Articles

Scroll down

to load the next article