Alibaba logo

Qwen3.8 Max

Multimodal
Alibaba

Qwen3.8 Max is Qwen's 2.4T-parameter multimodal flagship released in August 2026. It is a state-of-the-art performer on research replication (PaperBench 0.93) and long-context retrieval (MRCR v2 8-needle 0.93). It is served by DeepInfra, Fireworks, Novita and Together.

Key Specifications

Parameters
2.4T
Context
1.0M
Release Date
August 2, 2026
Average Score
71.0%

Timeline

Key dates in the model's history
Announcement
August 2, 2026
Last Update
August 27, 2026
Today
September 10, 2026

Technical Specifications

Parameters
2.4T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$6.25
Max Input Tokens
1.0M
Max Output Tokens
131.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
Self-reported
92.6%

Other Tests

Specialized benchmarks
PaperBench
BasicAgent Code-Dev, average of 3 runsSelf-reported
93.0%
MRCR v2 (8-needle)
256K context; 8 needlesSelf-reported
93.0%
Agents' Last Exam
ScoreSelf-reported
52.4%
AndroidBench
Self-reported
75.1%
AndroidWorld
Self-reported
85.3%
AutomationBench
Pass@1Self-reported
27.3%
CoWorkBench
Self-reported
74.8%
DeepSWE 1.1
Best result across Claude Code and mini-SWE-agentSelf-reported
56.6%
ERQA
Self-reported
77.8%
FrontierSWE
Claude Code, MEAN@5Self-reported
73.5%
HealthBench
Self-reported
60.2%
Humanity's Last Exam
Self-reported
43.6%
Humanity's Last Exam (with tools, text-only)
With toolsSelf-reported
56.2%
IFBench
Self-reported
82.8%
Job Bench
Self-reported
53.4%
LongBench v2
Self-reported
66.3%
LVBench
Self-reported
81.8%
MLS-Bench Lite
Claude Code, 5-hour timeoutSelf-reported
41.0%
MMMU-Pro
Qwen in-house evaluationSelf-reported
82.3%
MobileWorld
Self-reported
77.8%
NL2Repo
Claude Code harnessSelf-reported
55.9%
OneMillion Bench
Expert scoreSelf-reported
52.5%
OSWorld-Verified
Self-reported
86.1%
PerceptionBench
Qwen in-house evaluationSelf-reported
63.5%
PLawBench
Self-reported
73.2%
PRBench-Finance
Finance subsetSelf-reported
58.3%
PRBench-Legal
Legal subsetSelf-reported
57.6%
QwenQoderBench
Self-reported
58.4%
QwenReactBench
BT/Elo ratingSelf-reported
86.2%
QwenSVG
BT/Elo ratingSelf-reported
85.7%
QwenSWEBench
Self-reported
80.7%
RealWorldQA
Self-reported
88.0%
ScreenSpot Pro
Qwen in-house evaluationSelf-reported
84.5%
SkillsBench
SkillsBench v1.1, OpenCode, average of 3 runsSelf-reported
70.2%
SWE-Bench Pro
Claude Code, 256K contextSelf-reported
67.7%
Terminal-Bench 2.1
Claude Code, avg@10, 5-hour timeoutSelf-reported
86.6%
Toolathlon
Verified split; Pass@1Self-reported
72.5%
VideoMME w sub.
Subtitles enabledSelf-reported
90.4%
Vision2Web
Claude Code, GPT-5.4 judgeSelf-reported
69.0%
WideSearch
Self-reported
81.9%
Workspace Bench
Self-reported
67.7%

License & Metadata

License
proprietary
Announcement Date
August 2, 2026
Last Updated
August 27, 2026

Compare Qwen3.8 Max

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.