MiniCPM-SALA
MiniCPM-SALA (Sparse Attention and Linear Attention) is a 9B hybrid model built from a MiniCPM-4.0 checkpoint via continual training (~2T tokens, 25% of training-from-scratch cost). It interleaves 25% InfLLM-V2 sparse attention and 75% Lightning Attention layers, achieving up to 3.5x inference speed over dense baselines at 256K tokens. With HyPE (Hybrid Positional Encoding) and NoPE in sparse layers, the model extrapolates to 2048K tokens despite a 520K training length, enabling 1M-token inference on consumer GPUs like the RTX 5090.
Key Specifications
Timeline
Technical Specifications
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Other Tests
License & Metadata
Similar Models
All ModelsLFM2.5-2.6B
Liquid AI
Llama 3.1 8B Instruct
Meta
IBM Granite 4.0 Tiny Preview
IBM
Gemma 3 1B
DeepSeek R1 Distill Llama 8B
DeepSeek
Ministral 8B Instruct
Mistral AI
Llama 3.2 3B Instruct
Meta
DeepSeek R1 Distill Qwen 1.5B
DeepSeek
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.