Key Specifications
Parameters
-
Context
1.0M
Release Date
May 19, 2026
Average Score
73.6%
Timeline
Key dates in the model's history
Announcement
May 19, 2026
Last Update
August 27, 2026
Today
September 10, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$7.50
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
• Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Diamond • Self-reported
Other Tests
Specialized benchmarks
HMMT Feb 26
Self-reported by the model provider • Self-reported
Kernel Bench L3
Fraction of problems faster than torch.compile • Self-reported
BFCL-V4
• Self-reported
Claw-Eval
Average • Self-reported
CoWorkBench
• Self-reported
CritPT
• Self-reported
Finance Agent v2
• Verified
Global PIQA
• Self-reported
Humanity's Last Exam
• Self-reported
IFBench
• Self-reported
IFEval
Strict prompt • Self-reported
IMO-AnswerBench
• Self-reported
Include
• Self-reported
LiveBench
2026-01-08 • Verified
LiveCodeBench v6
• Self-reported
MathArena Apex
• Self-reported
MAXIFE
Accuracy on English and multilingual prompts across 23 settings • Self-reported
MCP Atlas
Public set, Gemini 2.5 Pro judger • Self-reported
MCP-Mark
GitHub MCP v0.30.3 • Self-reported
MMLU-Pro
• Self-reported
MMLU-ProX
Average accuracy across 29 languages • Self-reported
MMLU-Redux
• Self-reported
MMMLU
• Self-reported
MRCR 128K (8-needle)
MRCR-v2 128K context subset with 8 needles • Self-reported
NL2Repo
Claude Code scaffold • Self-reported
NOVA-63
• Self-reported
PolyMATH
• Self-reported
QwenSVG
Internal SVG generation benchmark, BT/Elo rating • Self-reported
QwenWebBench
QwenWebDev internal front-end code generation benchmark, BT/Elo rating • Self-reported
QwenWorldBench
Open-ended 5-dimensional rubric judge grounded in real-environment feedback • Self-reported
SciCode
• Self-reported
SkillsBench
OpenCode, 78 tasks, avg of 5 runs • Self-reported
SpreadSheetBench-v1
• Self-reported
SuperGPQA
• Self-reported
SWE-bench Multilingual
• Self-reported
SWE-Bench Pro
• Self-reported
Terminal-Bench 2.0
Terminus-2 harness • Self-reported
VITA-Bench
Average subdomain score • Self-reported
WMT24++
Average across 55 languages via XCOMET-XXL • Self-reported
ZClawBench
QwenClawBench real-user-distribution Claw agent benchmark • Self-reported
License & Metadata
License
proprietary
Announcement Date
May 19, 2026
Last Updated
August 27, 2026
Similar Models
All ModelsQwen3 Max
Alibaba
Best score:0.6 (GPQA)
Released:Dec 2025
Qwen3.5 122B A10B
Alibaba
122.0B
Best score:0.9 (GPQA)
Released:Mar 2026
Qwen3.5 35B A3B
Alibaba
35.0B
Best score:0.8 (GPQA)
Released:Mar 2026
Qwen3 Max Thinking
Alibaba
1.0T
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$1.20/1M tokens
Qwen3.5 27B
Alibaba
27.0B
Best score:0.9 (GPQA)
Released:Mar 2026
Qwen3.6 Plus
Alibaba
MM
Best score:0.9 (GPQA)
Released:Mar 2026
Price:$0.50/1M tokens
Qwen3.7-Plus
Alibaba
MM
Best score:0.9 (GPQA)
Released:May 2026
Price:$0.32/1M tokens
MiMo-V2-Pro
Xiaomi
Best score:1.0 (TAU)
Released:Mar 2026
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.