MiniCPM-SALA

OpenBMB

MiniCPM-SALA (Sparse Attention and Linear Attention) is a 9B hybrid model built from a MiniCPM-4.0 checkpoint via continual training (~2T tokens, 25% of training-from-scratch cost). It interleaves 25% InfLLM-V2 sparse attention and 75% Lightning Attention layers, achieving up to 3.5x inference speed over dense baselines at 256K tokens. With HyPE (Hybrid Positional Encoding) and NoPE in sparse layers, the model extrapolates to 2048K tokens despite a 520K training length, enabling 1M-token inference on consumer GPUs like the RTX 5090.

Key Specifications

Parameters
9.5B
Context
-
Release Date
February 11, 2026
Average Score
60.4%

Timeline

Key dates in the model's history
Announcement
February 11, 2026
Last Update
September 10, 2026
Today
September 19, 2026

Technical Specifications

Parameters
9.5B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
HumanEval
Self-reported
95.1%
MBPP
Self-reported
89.1%

Other Tests

Specialized benchmarks
AIME 2024
Self-reported
83.8%
AIME 2025
Self-reported
78.3%
BBH
Self-reported
81.5%
CMMLU
Self-reported
81.5%
IFEval
Self-reported
76.3%
LiveCodeBench v5
Self-reported
60.5%
LiveCodeBench v6
Self-reported
52.0%
MMLU-Pro
Self-reported
67.0%
MRCR 128K (2-needle)
Self-reported
28.6%
MRCR 128K (4-needle)
Self-reported
19.6%
MRCR 128K (8-needle)
Self-reported
10.1%
MRCR 64K (2-needle)
Self-reported
29.8%
MRCR 64K (4-needle)
Self-reported
20.6%
MRCR 64K (8-needle)
Self-reported
16.6%
NoLiMa 128K
Self-reported
23.9%
NoLiMa 32K
Self-reported
54.5%
NoLiMa 64K
Self-reported
43.0%
RULER 1000K
Self-reported
86.3%
RULER 128k
Self-reported
89.4%
RULER 2048K
Self-reported
81.6%
RULER 512K
Self-reported
87.1%
RULER 64k
Self-reported
92.7%

License & Metadata

License
apache_2_0
Announcement Date
February 11, 2026
Last Updated
September 10, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.