Alibaba logo

Qwen3.7 Max

Alibaba

Qwen3.7 Max is a proprietary flagship model from the Qwen team, released in May 2026. It leads Qwen's internal front-end generation benchmarks and tops HMMT Feb 26 with a 0.97 score, with strong agentic tool-use results on Kernel Bench L3. It is available via Novita and Together.

Key Specifications

Parameters
-
Context
1.0M
Release Date
May 19, 2026
Average Score
73.6%

Timeline

Key dates in the model's history
Announcement
May 19, 2026
Last Update
August 27, 2026
Today
September 10, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$7.50
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Self-reported
80.4%

Reasoning

Logical reasoning and analysis
GPQA
DiamondSelf-reported
92.4%

Other Tests

Specialized benchmarks
HMMT Feb 26
Self-reported by the model providerSelf-reported
97.0%
Kernel Bench L3
Fraction of problems faster than torch.compileSelf-reported
96.0%
BFCL-V4
Self-reported
75.0%
Claw-Eval
AverageSelf-reported
65.2%
CoWorkBench
Self-reported
67.2%
CritPT
Self-reported
11.4%
Finance Agent v2
Verified
48.4%
Global PIQA
Self-reported
91.4%
Humanity's Last Exam
Self-reported
41.4%
IFBench
Self-reported
79.1%
IFEval
Strict promptSelf-reported
94.3%
IMO-AnswerBench
Self-reported
90.0%
Include
Self-reported
86.2%
LiveBench
2026-01-08Verified
74.3%
LiveCodeBench v6
Self-reported
91.6%
MathArena Apex
Self-reported
44.5%
MAXIFE
Accuracy on English and multilingual prompts across 23 settingsSelf-reported
89.2%
MCP Atlas
Public set, Gemini 2.5 Pro judgerSelf-reported
76.4%
MCP-Mark
GitHub MCP v0.30.3Self-reported
60.8%
MMLU-Pro
Self-reported
89.6%
MMLU-ProX
Average accuracy across 29 languagesSelf-reported
87.0%
MMLU-Redux
Self-reported
95.0%
MMMLU
Self-reported
90.3%
MRCR 128K (8-needle)
MRCR-v2 128K context subset with 8 needlesSelf-reported
90.4%
NL2Repo
Claude Code scaffoldSelf-reported
47.2%
NOVA-63
Self-reported
59.0%
PolyMATH
Self-reported
86.5%
QwenSVG
Internal SVG generation benchmark, BT/Elo ratingSelf-reported
80.4%
QwenWebBench
QwenWebDev internal front-end code generation benchmark, BT/Elo ratingSelf-reported
78.4%
QwenWorldBench
Open-ended 5-dimensional rubric judge grounded in real-environment feedbackSelf-reported
57.3%
SciCode
Self-reported
53.5%
SkillsBench
OpenCode, 78 tasks, avg of 5 runsSelf-reported
59.2%
SpreadSheetBench-v1
Self-reported
87.0%
SuperGPQA
Self-reported
73.6%
SWE-bench Multilingual
Self-reported
78.3%
SWE-Bench Pro
Self-reported
60.6%
Terminal-Bench 2.0
Terminus-2 harnessSelf-reported
69.7%
VITA-Bench
Average subdomain scoreSelf-reported
47.9%
WMT24++
Average across 55 languages via XCOMET-XXLSelf-reported
85.8%
ZClawBench
QwenClawBench real-user-distribution Claw agent benchmarkSelf-reported
64.3%

License & Metadata

License
proprietary
Announcement Date
May 19, 2026
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.