Alibaba logo

Qwen3 Max Thinking

Alibaba

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Key Specifications

Parameters
1.0T
Context
256.0K
Release Date
February 13, 2026
Average Score
64.2%

Timeline

Key dates in the model's history
Announcement
February 13, 2026
Last Update
September 10, 2026
Today
September 20, 2026

Technical Specifications

Parameters
1.0T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.20
Output (per 1M tokens)
$6.00
Max Input Tokens
256.0K
Max Output Tokens
256.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Publisher comparison-table evaluation; scaffolding and agent setup not specified.Self-reported
75.3%

Reasoning

Logical reasoning and analysis
GPQA
Publisher comparison-table evaluation; setup not further specified.Self-reported
87.4%

Other Tests

Specialized benchmarks
AA-LCR
Publisher comparison-table evaluation; setup not further specified.Self-reported
68.7%
AIME 2026
Publisher comparison-table evaluation on AIME26; setup not further specified.Self-reported
93.3%
BFCL-V4
Publisher comparison-table evaluation; setup not further specified.Self-reported
67.7%
BrowseComp
Simple context-folding with a 256k context window.Self-reported
53.9%
BrowseComp-zh
Publisher comparison-table evaluation; setup not further specified.Self-reported
60.9%
C-Eval
Publisher comparison-table evaluation; setup not further specified.Self-reported
93.7%
DeepPlanning
Publisher comparison-table evaluation; setup not further specified.Self-reported
28.7%
Global PIQA
Publisher comparison-table evaluation; setup not further specified.Self-reported
86.0%
HLE-Verified
Verified and revised HLE subset with component-wise verification and a fine-grained error taxonomy.Self-reported
37.6%
Humanity's Last Exam (no tools, text-only)
Text-only HLE subset without tools.Self-reported
30.2%
Humanity's Last Exam (with tools, text-only)
Text-only HLE subset with tool use enabled.Self-reported
49.8%
IFBench
Publisher comparison-table evaluation; setup not further specified.Self-reported
70.9%
IFEval
Publisher comparison-table evaluation; setup not further specified.Self-reported
93.4%
IMO-AnswerBench
Publisher comparison-table evaluation; setup not further specified.Self-reported
83.9%
LiveCodeBench v6
Publisher comparison-table evaluation; setup not further specified.Self-reported
85.9%
LongBench v2
Publisher comparison-table evaluation; setup not further specified.Self-reported
60.6%
MAXIFE
Accuracy on English plus multilingual original prompts across 23 settings.Self-reported
84.0%
MCP-Mark
GitHub MCP server v0.30.3 from api.githubcopilot.com; Playwright responses truncated at 32k tokens.Self-reported
33.5%
MMLU-Pro
Publisher comparison-table evaluation; setup not further specified.Self-reported
85.7%
MMLU-Redux
Publisher comparison-table evaluation; setup not further specified.Self-reported
92.8%
MMMLU
Publisher comparison-table evaluation; setup not further specified.Self-reported
84.4%
Multi-Challenge
Publisher comparison-table evaluation; setup not further specified.Self-reported
63.3%
NOVA-63
Publisher comparison-table evaluation; setup not further specified.Self-reported
54.2%
Seal-0
Publisher comparison-table evaluation; setup not further specified.Self-reported
46.9%
SecCodeBench
Publisher comparison-table evaluation; setup not further specified.Self-reported
57.5%
SuperGPQA
Publisher comparison-table evaluation; setup not further specified.Self-reported
67.3%
SWE-bench Multilingual
Publisher comparison-table evaluation; scaffolding and agent setup not specified.Self-reported
66.7%
Terminal-Bench 2.0
Publisher comparison-table evaluation; agent harness and setup not specified.Self-reported
22.5%
Toolathlon
Publisher comparison-table evaluation; setup not further specified.Self-reported
18.8%
VITA-Bench
Publisher comparison-table evaluation; setup not further specified.Self-reported
40.9%
WideSearch
256k context window without context management.Self-reported
57.9%
WMT24++
Average across 55 languages using XCOMET-XXL on the harder rebalanced WMT24++ subset.Self-reported
77.6%

License & Metadata

License
proprietary
Announcement Date
February 13, 2026
Last Updated
September 10, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.