Key Specifications
Parameters
2.4T
Context
1.0M
Release Date
August 2, 2026
Average Score
71.0%
Timeline
Key dates in the model's history
Announcement
August 2, 2026
Last Update
August 27, 2026
Today
September 10, 2026
Technical Specifications
Parameters
2.4T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$6.25
Max Input Tokens
1.0M
Max Output Tokens
131.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Reasoning
Logical reasoning and analysis
GPQA
• Self-reported
Other Tests
Specialized benchmarks
PaperBench
BasicAgent Code-Dev, average of 3 runs • Self-reported
MRCR v2 (8-needle)
256K context; 8 needles • Self-reported
Agents' Last Exam
Score • Self-reported
AndroidBench
• Self-reported
AndroidWorld
• Self-reported
AutomationBench
Pass@1 • Self-reported
CoWorkBench
• Self-reported
DeepSWE 1.1
Best result across Claude Code and mini-SWE-agent • Self-reported
ERQA
• Self-reported
FrontierSWE
Claude Code, MEAN@5 • Self-reported
HealthBench
• Self-reported
Humanity's Last Exam
• Self-reported
Humanity's Last Exam (with tools, text-only)
With tools • Self-reported
IFBench
• Self-reported
Job Bench
• Self-reported
LongBench v2
• Self-reported
LVBench
• Self-reported
MLS-Bench Lite
Claude Code, 5-hour timeout • Self-reported
MMMU-Pro
Qwen in-house evaluation • Self-reported
MobileWorld
• Self-reported
NL2Repo
Claude Code harness • Self-reported
OneMillion Bench
Expert score • Self-reported
OSWorld-Verified
• Self-reported
PerceptionBench
Qwen in-house evaluation • Self-reported
PLawBench
• Self-reported
PRBench-Finance
Finance subset • Self-reported
PRBench-Legal
Legal subset • Self-reported
QwenQoderBench
• Self-reported
QwenReactBench
BT/Elo rating • Self-reported
QwenSVG
BT/Elo rating • Self-reported
QwenSWEBench
• Self-reported
RealWorldQA
• Self-reported
ScreenSpot Pro
Qwen in-house evaluation • Self-reported
SkillsBench
SkillsBench v1.1, OpenCode, average of 3 runs • Self-reported
SWE-Bench Pro
Claude Code, 256K context • Self-reported
Terminal-Bench 2.1
Claude Code, avg@10, 5-hour timeout • Self-reported
Toolathlon
Verified split; Pass@1 • Self-reported
VideoMME w sub.
Subtitles enabled • Self-reported
Vision2Web
Claude Code, GPT-5.4 judge • Self-reported
WideSearch
• Self-reported
Workspace Bench
• Self-reported
License & Metadata
License
proprietary
Announcement Date
August 2, 2026
Last Updated
August 27, 2026
Compare Qwen3.8 Max
All comparisonsSimilar Models
All ModelsQwen3.5-397B-A17B
Alibaba
MM397.0B
Best score:0.9 (GPQA)
Released:Feb 2026
Qwen3.8 Flash
Alibaba
MM125.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.15/1M tokens
Qwen3.8 Flash Next
Alibaba
MM180.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.16/1M tokens
Qwen3.6 Plus
Alibaba
MM
Best score:0.9 (GPQA)
Released:Mar 2026
Price:$0.50/1M tokens
Qwen3.8-27B
Alibaba
MM27.8B
Best score:0.9 (GPQA)
Released:Aug 2026
Qwen3.6-27B
Alibaba
MM27.8B
Best score:0.9 (GPQA)
Released:Apr 2026
Price:$0.60/1M tokens
Qwen3.7-Plus
Alibaba
MM
Best score:0.9 (GPQA)
Released:May 2026
Price:$0.32/1M tokens
Qwen3.6-35B-A3B
Alibaba
MM35.0B
Best score:0.9 (GPQA)
Released:Apr 2026
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.