All News
ifmk2-horizonmbzuaiopen-modelslocal-inferencellama-cpp

K2 Horizon 7B: a small model claims a 27B row, and the 'GPU poor' noticed

MBZUAI's IFM shipped a six-model K2 Horizon fleet with open recipes, data and logs. A 7B member claims 27B-class coding, and local runners are testing that.

Vlad MakarovVlad Makarovreviewed and published
2 min read
K2 Horizon 7B: a small model claims a 27B row, and the 'GPU poor' noticed

The Institute of Foundation Models at Abu Dhabi's MBZUAI shipped the K2 Horizon fleet on September 3: six Apache 2.0 models, from a 0.9B up to a 375B-A23B sparse mixture-of-experts, each pretrained on roughly 20 trillion tokens and each released with training code, data or data-construction recipes, intermediate checkpoints and fine-grained logs. Eleven days later a Reddit post handed the fleet's 7B member a nickname it never asked for.

The post that named it

On September 14, r/LocalLLaMA user u/Uncle___Marty posted a screenshot titled "For the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index." Read that as one poster's interpretation of a chart, not a published result: no Artificial Analysis row for K2 Horizon has been verified. The poster added that from initial testing the model "seems pretty solid so far," citing a llama.cpp CUDA build. The thread drew roughly 525 points and about 50 comments.

The framing is economic more than sentimental. RTX 5090s have been disappearing from US first-party retail at street prices from about $4,400 to a $10,850 bundle, and builders have been routing around that with used Epyc servers. At those prices a 7B is the cheapest road to a capable local model.

What is actually new here

A 7B with long context is not the story, and Qwen has already pushed 27B-class models onto local hardware. The GGUF card lists a native 524,288-token context. What K2 Horizon adds is the paper trail: architectures, mixture compositions, training code, configurations and checkpoints pulled from the run.

IFM's 7B card reports SWE-bench Verified at 70.6 versus 50.8 for Qwen3.5-9B and 30.6 for Gemma 4-12B, and Terminal-Bench 2.1 at 39.1 against 29.2 and 27.3. All of it is vendor-run, and the table carries its own counterexample: on SciCode, Gemma 4-12B scores 38.2 to the K2 model's 31.6. The card also notes that its BrowseComp comparison numbers may come from different harnesses.

The phone claim meets a pending pull request

MBZUAI's press release calls K2 Horizon "the largest fully open model release in AI history" and says the 7B is small enough to run on a phone. What shipped complicates that. The GGUF repository states that the files require a llama.cpp build with K2 Horizon architecture support, that the upstream pull request is still in progress, and points to a vendor fork branch. Its tensors stay in BF16, around 9B parameters and 18 GB. The phone story rests on quantization support the fleet advertises, not on anything a reader can run tonight.

What would settle the ranking

An Artificial Analysis row measured by the runner, plus independent reproductions of SWE-bench Verified and BrowseComp on the standard harnesses, would decide whether a 7B genuinely sits between two 27B-class rivals.

Related Articles

Scroll down

to load the next article