Key Specifications
Parameters
2.8T
Context
1.0M
Release Date
July 16, 2026
Average Score
67.8%
Timeline
Key dates in the model's history
Announcement
July 16, 2026
Last Update
August 27, 2026
Today
September 8, 2026
Technical Specifications
Parameters
2.8T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$3.00
Output (per 1M tokens)
$15.00
Max Input Tokens
1.0M
Max Output Tokens
1.0M
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Reasoning
Logical reasoning and analysis
GPQA
GPQA-Diamond; max reasoning effort • Self-reported
Other Tests
Specialized benchmarks
MathVision
With Python; average of 3 runs • Self-reported
DeepSearchQA
F1 score; max reasoning effort • Self-reported
AA-Briefcase
Elo score; Artificial Analysis • Verified
APEX-Agents
Max reasoning effort • Self-reported
AutomationBench
600-task public subset; official setup • Self-reported
BabyVision
With Python; average of 3 runs • Self-reported
BrowseComp
Context compaction at 300K tokens; max reasoning effort • Self-reported
CharXiv-R
Reasoning questions with Python; average of 3 runs • Self-reported
DECK-Bench
Internal benchmark; max reasoning effort • Self-reported
DeepSWE
KimiCode harness; max reasoning effort • Self-reported
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 69% ± 5%. • Verified
FrontierSWE
Dominance score; KimiCode harness; max reasoning effort • Self-reported
Humanity's Last Exam
HLE-Full with tools; max reasoning effort • Self-reported
Job Bench
Max reasoning effort • Self-reported
Kimi Code Bench v2
Internal benchmark; KimiCode and Claude Code harnesses; max reasoning effort • Self-reported
MCP Atlas
500-task public subset; 100-turn limit; Gemini 3.1 Pro judge • Self-reported
MLS-Bench Lite
KimiCode harness; max reasoning effort • Self-reported
MMMU-Pro
Official protocol; average of 3 runs • Self-reported
MMMU-Pro (with tools)
Python tools; official protocol; average of 3 runs • Self-reported
OfficeQA Pro
Claude Code harness; max reasoning effort • Self-reported
OmniDocBench
Average of 3 runs • Self-reported
PerceptionBench
Internal benchmark; average of 3 runs • Self-reported
PostTrainBench
Official Harbor implementation; Claude Code harness; max reasoning effort; average of 3 runs • Self-reported
Program Bench
KimiCode harness; max reasoning effort • Self-reported
SpreadsheetBench 2
Claude Code harness; max reasoning effort • Self-reported
SWE-Marathon
Claude Code harness; max reasoning effort • Self-reported
Terminal-Bench 2.1
KimiCode harness; max reasoning effort • Self-reported
Toolathlon
Toolathlon-Verified; max reasoning effort • Self-reported
WorldVQA
ForceAnswer; average of 3 runs • Self-reported
ZEROBench
Main split with Python; pass@5; 5 runs • Self-reported
License & Metadata
License
proprietary
Announcement Date
July 16, 2026
Last Updated
August 27, 2026
Compare Kimi K3
All comparisonsSimilar Models
All ModelsKimi K2.6
Moonshot AI
MM1.0T
Best score:0.9 (GPQA)
Released:Apr 2026
Price:$1.20/1M tokens
Kimi K2.5
Moonshot AI
MM1.0T
Best score:0.9 (GPQA)
Released:Jan 2026
Kimi-k1.5
Moonshot AI
MM
Best score:0.9 (MMLU)
Released:Jan 2025
Step-3.5-Flash
StepFun
MM196.0B
Best score:0.9 (TAU)
Released:Feb 2026
Price:$0.10/1M tokens
Kimi K2 0905
Moonshot AI
1.0T
Best score:0.9 (HumanEval)
Released:Sep 2025
Price:$0.60/1M tokens
Command A+
Cohere
MM218.0B
Best score:0.8 (TAU)
Released:May 2026
Inkling-Small
Thinking Machines Lab
MM276.0B
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$0.30/1M tokens
Qwen3.8 Flash
Alibaba
MM125.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.15/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.