Moonshot AI logo

Kimi K3

Multimodal
Moonshot AI

Kimi K3 is Moonshot AI's flagship 2.8T-parameter multimodal MoE model released in July 2026. It excels at document understanding and multimodal math, scoring 0.98 on MathVision and 0.94 on GPQA-Diamond. It is served by Fireworks, Moonshot AI, Novita and Together.

Key Specifications

Parameters
2.8T
Context
1.0M
Release Date
July 16, 2026
Average Score
67.8%

Timeline

Key dates in the model's history
Announcement
July 16, 2026
Last Update
August 27, 2026
Today
September 8, 2026

Technical Specifications

Parameters
2.8T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$3.00
Output (per 1M tokens)
$15.00
Max Input Tokens
1.0M
Max Output Tokens
1.0M
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
GPQA-Diamond; max reasoning effortSelf-reported
94.0%

Other Tests

Specialized benchmarks
MathVision
With Python; average of 3 runsSelf-reported
98.0%
DeepSearchQA
F1 score; max reasoning effortSelf-reported
95.0%
AA-Briefcase
Elo score; Artificial AnalysisVerified
51.6%
APEX-Agents
Max reasoning effortSelf-reported
37.6%
AutomationBench
600-task public subset; official setupSelf-reported
30.8%
BabyVision
With Python; average of 3 runsSelf-reported
85.7%
BrowseComp
Context compaction at 300K tokens; max reasoning effortSelf-reported
91.2%
CharXiv-R
Reasoning questions with Python; average of 3 runsSelf-reported
91.3%
DECK-Bench
Internal benchmark; max reasoning effortSelf-reported
73.5%
DeepSWE
KimiCode harness; max reasoning effortSelf-reported
67.5%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 69% ± 5%.Verified
69.0%
FrontierSWE
Dominance score; KimiCode harness; max reasoning effortSelf-reported
81.2%
Humanity's Last Exam
HLE-Full with tools; max reasoning effortSelf-reported
56.0%
Job Bench
Max reasoning effortSelf-reported
52.9%
Kimi Code Bench v2
Internal benchmark; KimiCode and Claude Code harnesses; max reasoning effortSelf-reported
72.9%
MCP Atlas
500-task public subset; 100-turn limit; Gemini 3.1 Pro judgeSelf-reported
84.2%
MLS-Bench Lite
KimiCode harness; max reasoning effortSelf-reported
48.3%
MMMU-Pro
Official protocol; average of 3 runsSelf-reported
81.6%
MMMU-Pro (with tools)
Python tools; official protocol; average of 3 runsSelf-reported
83.4%
OfficeQA Pro
Claude Code harness; max reasoning effortSelf-reported
63.3%
OmniDocBench
Average of 3 runsSelf-reported
91.1%
PerceptionBench
Internal benchmark; average of 3 runsSelf-reported
58.5%
PostTrainBench
Official Harbor implementation; Claude Code harness; max reasoning effort; average of 3 runsSelf-reported
36.6%
Program Bench
KimiCode harness; max reasoning effortSelf-reported
77.8%
SpreadsheetBench 2
Claude Code harness; max reasoning effortSelf-reported
34.8%
SWE-Marathon
Claude Code harness; max reasoning effortSelf-reported
42.0%
Terminal-Bench 2.1
KimiCode harness; max reasoning effortSelf-reported
88.3%
Toolathlon
Toolathlon-Verified; max reasoning effortSelf-reported
73.2%
WorldVQA
ForceAnswer; average of 3 runsSelf-reported
51.0%
ZEROBench
Main split with Python; pass@5; 5 runsSelf-reported
41.0%

License & Metadata

License
proprietary
Announcement Date
July 16, 2026
Last Updated
August 27, 2026

Compare Kimi K3

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.