Qwen3.8 Flash vs Qwen3.8 Max: Specs & Benchmark Comparison

Qwen3.8 Flash and Qwen3.8 Max both come from Alibaba, and both were released in August 2026. Qwen3.8 Max is the larger model at roughly 2.4 trillion parameters, against 125 billion for Qwen3.8 Flash.

The two models share 15 published benchmarks. Qwen3.8 Max leads on 11 of them, Qwen3.8 Flash on 4. The widest gaps are on NL2Repo, where Qwen3.8 Max scores 55.9% against 48.1%; Humanity's Last Exam, where Qwen3.8 Max scores 43.6% against 35.9%. Averaged across everything we track, Qwen3.8 Flash sits at 68.5% and Qwen3.8 Max at 71.0%.

Qwen3.8 Flash is the cheaper API at $0.15 per million input tokens and $0.47 per million output tokens, roughly 17 times cheaper than Qwen3.8 Max at $2.50 and $6.25. Both accept a context window of about 1M tokens. Qwen3.8 Flash accepts text, images, and video as input, while Qwen3.8 Max accepts text and images. On tooling, only Qwen3.8 Max supports function calling and only Qwen3.8 Max offers structured output.

CharacteristicQwen3.8 FlashQwen3.8 Max
CompanyAlibabaAlibaba
Release DateAugust 26, 2026August 2, 2026
Parameters125B2.4T
MultimodalYesYes
Context (input)1.0M1.0M
Context (output)131K131K
Input Price / 1M$0.15$2.50
Output Price / 1M$0.47$6.25
Average Score68.5%71.0%
Benchmarks
NL2Repo48.1%55.9%
Humanity's Last Exam35.9%43.6%
ERQA72.3%77.8%
SWE-Bench Pro62.5%67.7%
LVBench76.6%81.8%
Vision2Web64.0%69.0%
Job Bench55.7%53.4%
DeepSWE 1.158.7%56.6%
IFBench81.3%82.8%
Agents' Last Exam51.2%52.4%
Toolathlon73.5%72.5%
CoWorkBench73.9%74.8%
GPQA91.7%92.6%
AndroidWorld84.5%85.3%
RealWorldQA88.5%88.0%

Visual Benchmark Comparison

Qwen3.8 Flash
Qwen3.8 Max
NL2Repo0.5 vs 0.6
0.5
0.6
Humanity's Last Exam0.4 vs 0.4
0.4
0.4
ERQA0.7 vs 0.8
0.7
0.8
SWE-Bench Pro0.6 vs 0.7
0.6
0.7
LVBench0.8 vs 0.8
0.8
0.8
Vision2Web0.6 vs 0.7
0.6
0.7
Job Bench0.6 vs 0.5
0.6
0.5
DeepSWE 1.10.6 vs 0.6
0.6
0.6
IFBench0.8 vs 0.8
0.8
0.8
Agents' Last Exam0.5 vs 0.5
0.5
0.5
Toolathlon0.7 vs 0.7
0.7
0.7
CoWorkBench0.7 vs 0.7
0.7
0.7
GPQA0.9 vs 0.9
0.9
0.9
AndroidWorld0.8 vs 0.9
0.8
0.9
RealWorldQA0.9 vs 0.9
0.9
0.9

Verdict

Qwen3.8 Flash leads in 3 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: Qwen3.8 Flash — 0.7, Qwen3.8 Max — 0.7.

API Cost

Qwen3.8 Flash is 14.1x cheaper: input $0.15/1M vs $2.50/1M tokens.

Context Window

Qwen3.8 Flash supports a larger context: 1M vs 1M tokens.

Recency

Qwen3.8 Flash is newer: released 8/26/2026 vs 8/2/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — Qwen3.8 Flash or Qwen3.8 Max?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — Qwen3.8 Flash or Qwen3.8 Max?
Qwen3.8 Flash is cheaper for input: $0.15 per 1M tokens vs $2.50.
Which has a larger context window — Qwen3.8 Flash or Qwen3.8 Max?
Qwen3.8 Flash supports a larger context: 1,048,576 tokens vs 1,000,000.

The Qwen3.8 Flash and Qwen3.8 Max comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the Qwen3.8 Flash or Qwen3.8 Max page. See also the complete list of AI model comparisons.