Inkling
MultimodalInkling is Thinking Machines Lab's 975B-parameter open-weights Mixture-of-Experts model, released under Apache 2.0 on 2026-07-21. It accepts text, image, and audio input and generates text, with native reasoning and variable thinking effort (the published figures use effort=0.99). The vendor model card reports 97.1% on AIME 2026, 77.6% on SWE-Bench Verified, 73.5% on MMMU-Pro, 77.2% on audio MMAU, and 91.4% on the VoiceBench average, positioning it as the flagship sibling of the smaller Inkling-Small. It is served by DeepInfra with a 512K-token context window at $0.95 / $4.05 per 1M input/output tokens.
Key Specifications
Timeline
Technical Specifications
Pricing & Availability
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Other Tests
License & Metadata
Compare Inkling
All comparisonsSimilar Models
All ModelsInkling-Small
Thinking Machines Lab
MiniMax M2.5
MiniMax
MiniMax M3
MiniMax
MiMo-V2.5
Xiaomi
MAI-Code-1.1-Flash
Microsoft
Mistral Medium 3.5
Mistral AI
GLM-5.3 Flash
Zhipu AI
Kimi K2.7 Code
Moonshot AI
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.