All News
qwenlocal-llmresearchllama-cppn-gram

Qwengram-0.8B reattaches Qwen's n-gram memory to a frozen tiny model

An independent researcher attached Qwen3.8-Flash-Next's frozen n-gram memory to a frozen Qwen3.5-0.8B backbone, cutting validation perplexity by about 5%.

Vlad MakarovVlad Makarovreviewed and published
3 min read
Qwengram-0.8B reattaches Qwen's n-gram memory to a frozen tiny model

An independent researcher publishing as Ninnix96 has released Qwengram-0.8B, an experiment that lifts the n-gram memory out of Qwen's Qwen3.8-Flash-Next and attaches it to a frozen Qwen3.5-0.8B backbone. Only a small reader trains. Frozen validation perplexity falls 5.048%, from 18.2759 to 17.3534.

A 51B memory table bolted onto a frozen 0.8B model

It tests whether the n-gram embedding table Qwen bet on in August is portable. Qwen3.8-Flash-Next shipped with a 125B-parameter main model plus roughly 51B extra n-gram PLE embedding parameters, 6B activated per token. Qwengram keeps that memory frozen and wires it into a far smaller backbone through a trainable reader (R=1) at layers 3 and 9, behind a dynamic token-level gate. The reader saw 15,000,064 tokens, 749,568 of them calibration, trained mostly on free Kaggle GPUs. No backbone fine-tuning.

The numbers:

  • Frozen validation perplexity: 18.2759 to 17.3534 (-5.048%); NLL 2.905585 to 2.853786
  • Canonical checkpoint: LAMBADA-1000 45.6% accuracy, HellaSwag-1000 40.0%, five-domain mean NLL 2.384680
  • GGUF (excl. sidecar): Q4_K_M 584 MB, Q6_K 688 MB, Q8_0 876 MB, BF16 1.60 GB; Q8_0 keeps 99.1% of the gain
  • Runtime: a ~32 GB Q4_1 quantized PLE sidecar (memory-mapped) plus a pinned llama.cpp fork; stock upstream cannot run the reader

The controls matter more than the headline. The real pretrained memory beat both random-memory and permuted-memory baselines; with the same budget, learned token placement beat shuffled placement in every domain; and the gate's strength varies substantially across tokens rather than acting like a learned constant.

Five percent, and the asterisks attached to it

The author is careful about the number. "This is a language-model validation result, not a claim of 5% higher benchmark accuracy," the write-up states, and the card concedes that "Benchmark accuracy differences were not established." A 5% perplexity gain on a 0.8B model demonstrates a mechanism, not a capability jump: these are language-modelling metrics, not task scores. The friction is real: the 48.7 GiB PLE source is not packed into the released files, so running it needs a third-party ~32 GB sidecar and a pinned fork, a size ceiling that echoes this month's 16 GB VRAM discussion. The 20M-token endpoint cut aggregate loss further but regressed on math, so 15M stayed the balanced checkpoint; LAMBADA was hurt by strong fixed late-layer injection, which the dynamic gate only partly recovered.

Three stars, 321 downloads, one open question

What's next is mostly community. The model and GGUFs drew 321 downloads in the last month; the code is MIT-licensed and credits a paper on target-side reader adaptation. Reproducibility is partial: three stars, two forks, some private evaluation artifacts, and a Reddit thread at roughly 350 points and 95 comments. The open question is whether other frozen backbones can borrow Qwen's n-gram memory, and whether anyone with a real evaluation budget tests it on tasks rather than perplexity.

Related Articles

Scroll down

to load the next article