JetBrains Mellum2.1: a new training recipe in an unchanged 12B model
JetBrains' Mellum2.1 keeps the 12B Mellum2 architecture and moves reinforcement learning to the main training stage. We check what shipped and what did not.
JetBrains has released Mellum2.1, a 12B mixture-of-experts coding model whose architecture is identical to the Mellum2 build the company open-sourced in June 2026. Total parameters are the same 12B, active parameters the same roughly 2.5B, and the licence the same Apache 2.0. What changed is everything after pre-training. This is not a bigger model; it is a differently trained one, and the vendor's central claim is that reinforcement learning moved from a short final stage to the main part of the pipeline.
| Spec | Mellum2.1 |
|---|---|
| Total parameters | 12B |
| Active parameters | ~2.5B |
| Licence | Apache 2.0 |
| Announced | 8 October 2026 |
| Availability | Hugging Face |
A training recipe, sold as a release
JetBrains is explicit that it is selling a method, not a design. In its own framing, Mellum2 "was fast, but it couldn't work inside a repository at the level we wanted", and the answer was "a summer of reinforcement learning in real environments, with millions of sandboxed runs across thousands of environments" — per the announcement post — after which the model can explore a codebase, edit files and check its own changes. New RL tasks were added in math, competitive programming, science, tool use and software engineering, mixing open datasets with JetBrains-authored tasks. Every source was filtered before use, and the stated reason is revealing: "Open data often comes with broken tests, unverifiable answers, or tasks that are too easy or impossible for the model." That sentence is as much a comment on how public RL corpora arrive as a description of JetBrains' pipeline.
Making RL the main stage rather than a final polish also changes what the model is optimised for: not next-token likelihood on code, but sequences of actions inside a repository — reading, editing, running tests and revising. That is a harder objective and a less stable one, which is why the environment count, not the parameter count, is the number worth watching.
The comparison set is a claim, not a result
The Hugging Face collection ranks the model against Mellum2, Qwen3.5-9B and Gemma 4 E4B on one JetBrains-assembled evaluation setup. The named benchmarks are LiveCodeBench, AIME, GPQA, BFCL, IFEval and SWE-bench Verified; the vendor reports gains across all of them, with the largest improvement over Mellum2 in agentic coding, alongside smaller gains in coding, competitive programming, math, tool calling and general knowledge. The figures themselves exist only as charts, not as numbers in the text, so nothing here can be checked against an outside run. Rival numbers are assembled by the vendor too, which is normal for a launch and also why the comparison cannot stand alone. JetBrains does not publish the raw prompts, seeds or serving configuration behind the charts, so even the direction of travel is taken on trust. No independent evaluation shipped with the release. Read the charts as a direction of travel, not a league table.
Speed numbers come from an H200
Post-training did not touch the architecture, so Mellum2.1 is as fast as Mellum2, and multi-token prediction is meant to make it faster. JetBrains says that under heavy load the model serves "almost twice as many tokens as Qwen3.5-9B", and that MTP makes a single request "about 1.6 times faster". The unit is output tokens per second on a single H200 — a datacentre GPU. That is worth holding against the rest of the pitch. The gain rides on speculative decoding with an MTP head, and vLLM's MTP support landed days earlier, but the head Mellum2.1 needs for that speedup was not part of the release. MTP proposes several future tokens per step and a smaller draft pass verifies them; its benefit is largest when the server is under light load. When the head arrives, the effective speedup will depend on batch size and the draft acceptance rate, neither of which the vendor reports.
Availability: the hardware pitch is a promise
Availability today is Hugging Face, and the Apache 2.0 licence is genuinely permissive. What is not yet here is the local runtime story. GGUF builds for llama.cpp, Ollama and LM Studio, along with the MTP head for speculative decoding in vLLM, are "coming soon" in JetBrains' own wording. Until those land, "runs on your own hardware" describes an intent rather than an experience most users can have. A trend report on r/LocalLLaMA put the target at roughly 16 GB of VRAM with a Q6 MLX variant promised — a plan, not a shipped artifact. The gap is not unusual for an open release, since quantisations often trail base weights by days, but the whole selling point here is agentic work on hardware you own, and that is the part still pending.
Thin adoption for a load-bearing claim
A snapshot of the Hugging Face collection on 9-10 October showed modest uptake. The Thinking model sat at roughly 198 downloads and 76 likes; the GGUF build at roughly 1,200 downloads and 41 likes; the collection itself at 13 upvotes. JetBrains also positions the model as a general assistant for everyday questions and harder math and reasoning, and as private, self-hosted deployment that keeps code and data under the customer's control — a common enterprise framing that the thin local tooling undercuts in the near term. On Reddit, the r/LocalLLaMA thread "Mellum2.1 - a JetBrains Collection" drew roughly 123 points and 63 comments, where a JetBrains engineer answered questions directly and said the biggest gains were in agentic coding versus Mellum2. Download counts are a running total and a snapshot, not a verdict, but the shape is a working thread, not a breakout, and the numbers behind the central "gains" claim remain vendor-reported on JetBrains' own comparison set.
What would make the claim checkable
What is missing is not effort but evidence a buyer can reproduce. Independent runs of the same benchmarks on the same quantised weights would settle whether a 2.5B-active model does reliable multi-file repository work. Shipped GGUF and MTP support would show the speed claim survives outside an H200 and outside JetBrains' own serving stack. A reproducible agentic-coding evaluation, with the environment and the task set published alongside the numbers, would show the RL recipe does what the charts imply. Until one of those lands, the honest summary is that JetBrains has shipped a training method wrapped in a release — and the deployment story behind it is still a promise.


