Key Specifications
Parameters
-
Context
1.0M
Release Date
May 31, 2026
Average Score
70.3%
Timeline
Key dates in the model's history
Announcement
May 31, 2026
Last Update
August 29, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.32
Output (per 1M tokens)
$1.28
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
• Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Diamond • Self-reported
Other Tests
Specialized benchmarks
AndroidWorld
• Self-reported
Apex
• Self-reported
BabyVision
With CI • Self-reported
BC-VL
With search • Self-reported
BFCL-V4
• Self-reported
CharXiv-R
With CI • Self-reported
Claw-Eval
• Self-reported
ClawEval-MM
• Self-reported
CountQA
• Self-reported
CoWorkBench
• Self-reported
CritPT
• Self-reported
DeepPlanning
• Self-reported
ERQA
• Self-reported
Finance Agent v2
• Verified
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score: 10.2%. • Verified
Global PIQA
• Self-reported
HiPhO
• Self-reported
HMMT Feb 26
• Self-reported
Humanity's Last Exam
• Self-reported
IFBench
• Self-reported
IFEval
• Self-reported
IMO-AnswerBench
• Self-reported
Include
• Self-reported
LingoQA
• Self-reported
LiveCodeBench v6
• Self-reported
LVBench
• Self-reported
MathVision
• Self-reported
MAXIFE
• Self-reported
MCP Atlas
Public set • Self-reported
MCP-Mark
• Self-reported
MedXpertQA-MM
• Self-reported
MLVU
M-Avg • Self-reported
MMBC
With search • Self-reported
MMLU-Pro
• Self-reported
MMLU-ProX
• Self-reported
MMLU-Redux
• Self-reported
MMMLU
• Self-reported
MMMU-Pro
• Self-reported
MMSearch-Plus
With search • Self-reported
MRCR v2
128k • Self-reported
NL2Repo
• Self-reported
NOVA-63
• Self-reported
OCRBench_V2
Chinese • Self-reported
ODinW
13 datasets • Self-reported
OmniDocBench 1.5
• Self-reported
OSWorld-Verified
enable_thinking=False • Self-reported
PolyMATH
• Self-reported
QwenClawBench
• Self-reported
QwenWorldBench
• Self-reported
RealWorldQA
• Self-reported
SciCode
• Self-reported
ScreenSpot Pro
enable_thinking=False • Self-reported
SimpleVQA
With search • Self-reported
SkillsBench
• Self-reported
SpreadSheetBench-v1
• Self-reported
SuperGPQA
• Self-reported
SURDS
• Self-reported
SWE-bench Multilingual
• Self-reported
SWE-Bench Pro
• Self-reported
Terminal-Bench 2.0
Terminus-2 • Self-reported
TVBench
• Self-reported
Video-MME
With subtitles • Self-reported
VideoMMMU
• Self-reported
VisFactor
• Self-reported
VITA-Bench
• Self-reported
VLADBench
• Self-reported
WMT24++
• Self-reported
WorldVQA
With search • Self-reported
License & Metadata
License
proprietary
Announcement Date
May 31, 2026
Last Updated
August 29, 2026
Similar Models
All ModelsQwen3.6 Plus
Alibaba
MM
Released:Mar 2026
Price:$0.50/1M tokens
Qwen3.6-35B-A3B
Alibaba
MM35.0B
Best score:0.9 (GPQA)
Released:Apr 2026
Qwen3.8-27B
Alibaba
MM27.8B
Best score:0.9 (GPQA)
Released:Aug 2026
Qwen3.8 Flash Next
Alibaba
MM180.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.16/1M tokens
ERNIE 5.0
Baidu
MM
Best score:0.8 (GPQA)
Released:Jan 2025
Seed 2.0 Pro
ByteDance
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$0.50/1M tokens
Kimi-k1.5
Moonshot AI
MM
Best score:0.9 (MMLU)
Released:Jan 2025
Gemini 3.1 Pro
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$2.50/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.