Microsoft logo

MAI-Thinking-1

Microsoft

MAI-Thinking-1 is Microsoft AI's first in-house reasoning model: a 35B-active / ~1T-total sparse Mixture-of-Experts model trained from scratch without distillation from third-party models. Built with Microsoft's Hill-Climbing Machine pipeline on 30T tokens of clean licensed data, it goes toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro and scores 97.0% on AIME 2025. Supports a 256K context window, function calling, and developer instructions.

Key Specifications

Parameters
1.0T
Context
-
Release Date
June 2, 2026
Average Score
72.3%

Timeline

Key dates in the model's history
Announcement
June 2, 2026
Last Update
August 31, 2026

Technical Specifications

Parameters
1.0T
Training Tokens
30.0T tokens
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
TruthfulQA
Self-reported
88.0%

Programming

Programming skills tests
SWE-Bench Verified
256k contextSelf-reported
73.5%

Reasoning

Logical reasoning and analysis
GPQA
DiamondSelf-reported
84.2%

Other Tests

Specialized benchmarks
AdvancedIF
Rubric-levelSelf-reported
85.0%
AIME 2025
Self-reported
97.0%
AIME 2026
Self-reported
94.5%
AIR-Bench
Self-reported
88.0%
BFCL-v3
Self-reported
72.0%
CorpusQA
Self-reported
82.0%
CyberSecEval 4
Insecure code: AutocompleteSelf-reported
63.0%
GraphWalks
F1, ≤128k contextSelf-reported
90.0%
HealthBench Professional
Self-reported
35.0%
HMMT Feb 26
Self-reported
84.9%
IFBench
Self-reported
69.0%
LiveCodeBench v6
Self-reported
87.7%
LongBench v2
256k contextSelf-reported
61.0%
LongFact
Self-reported
98.0%
MedXpertQA
TextSelf-reported
43.0%
MMLU-Pro
Self-reported
85.0%
Multi-Challenge
Self-reported
53.0%
SimpleQA Verified
Self-reported
31.0%
SWE-Bench Pro
256k contextSelf-reported
52.8%
Terminal-Bench 2.0
256k contextSelf-reported
46.0%

License & Metadata

License
proprietary
Announcement Date
June 2, 2026
Last Updated
August 31, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.