Xiaomi logo

MiMo-V2-Omni

Multimodal
Xiaomi

MiMo-V2-Omni is Xiaomi's omni foundation model uniting frontier multimodal understanding with strong agentic capability. It fuses dedicated image, video, and audio encoders into a single shared backbone, processing all modalities simultaneously. Natively supports structured tool calling, function execution, and UI grounding. Supports over 10 hours of continuous audio understanding and 256K token context window. It is served by xiaomi with a 256K-token context window at $0.4 / $2 per 1M input/output tokens.

Key Specifications

Parameters
-
Context
262.0K
Release Date
March 18, 2026
Average Score
62.5%

Timeline

Key dates in the model's history
Announcement
March 18, 2026
Last Update
September 12, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.40
Output (per 1M tokens)
$2.00
Max Input Tokens
262.0K
Max Output Tokens
16.4K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
• Self-reported
74.8%

Other Tests

Specialized benchmarks
Claw-Eval
• Self-reported
54.8%
MM-BrowserComp
• Self-reported
52.0%
OmniGAIA
• Self-reported
49.8%
PinchBench
Average • Self-reported
81.2%

License & Metadata

License
proprietary
Announcement Date
March 18, 2026
Last Updated
September 12, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.