Moonshot AI logo

Kimi K2-Thinking-0905

Moonshot AI

Kimi K2 Thinking 0905 is the September 2025 update to the reasoning-focused variant of Kimi K2. It is a large-scale Mixture-of-Experts (MoE) language model by Moonshot AI with 1 trillion parameters and 32 billion active per forward pass. The thinking variant is specifically optimized for complex reasoning, mathematical proofs, and multi-step problem solving. This update enhances agentic coding, frontend development, and supports long-context inference up to 256K tokens.

Key Specifications

Parameters
1.0T
Context
262.1K
Release Date
September 4, 2025
Average Score
68.3%

Timeline

Key dates in the model's history
Announcement / Last Update
September 4, 2025
Today
September 11, 2026

Technical Specifications

Parameters
1.0T
Training Tokens
15.5T tokens
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.60
Output (per 1M tokens)
$2.40
Max Input Tokens
262.1K
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
w/ toolsSelf-reported
71.3%

Reasoning

Logical reasoning and analysis
GPQA
DiamondSelf-reported
84.5%

Other Tests

Specialized benchmarks
AIME 2025
Standard evaluationSelf-reported
100.0%
HMMT 2025
Standard evaluationSelf-reported
97.5%
MMLU-Redux
Without toolsSelf-reported
94.0%
MMLU-Pro
Without toolsSelf-reported
85.0%
BrowseComp
w/ toolsSelf-reported
60.2%
BrowseComp-zh
w/ toolsSelf-reported
62.3%
FinSearchComp-T3
w/ toolsSelf-reported
47.4%
FRAMES
w/ toolsSelf-reported
87.0%
HealthBench
no toolsSelf-reported
58.0%
Humanity's Last Exam
heavySelf-reported
51.0%
IMO-AnswerBench
no toolsSelf-reported
78.6%
LiveCodeBench v6
no toolsSelf-reported
83.1%
Multi-SWE-Bench
w/ toolsSelf-reported
41.9%
OJBench
cpp, no toolsSelf-reported
48.7%
SciCode
no toolsSelf-reported
44.8%
Seal-0
w/ toolsSelf-reported
56.3%
SWE-bench Multilingual
w/ toolsSelf-reported
61.1%
Terminal-Bench
w/ simulated tools (JSON)Self-reported
47.1%
WritingBench
Longform Writing, no toolsSelf-reported
73.8%

License & Metadata

License
mit
Announcement Date
September 4, 2025
Last Updated
September 4, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.