Qwen3-235B-A22B-Thinking-2507 vs Qwen3.8 Flash: Specs & Benchmark Comparison

Qwen3-235B-A22B-Thinking-2507 and Qwen3.8 Flash both come from Alibaba, Qwen3-235B-A22B-Thinking-2507 was released in July 2025, and Qwen3.8 Flash followed 13 months later in August 2026. Qwen3-235B-A22B-Thinking-2507 is the larger model at roughly 235 billion parameters, against 125 billion for Qwen3.8 Flash.

The two models share 3 published benchmarks. Qwen3.8 Flash leads on 3 of them. The widest gaps are on LiveCodeBench v6, where Qwen3.8 Flash scores 91.9% against 74.1%; Humanity's Last Exam, where Qwen3.8 Flash scores 35.9% against 18.2%. Averaged across everything we track, Qwen3-235B-A22B-Thinking-2507 sits at 69.2% and Qwen3.8 Flash at 68.5%.

Qwen3.8 Flash is the cheaper API at $0.15 per million input tokens and $0.47 per million output tokens, roughly 2 times cheaper than Qwen3-235B-A22B-Thinking-2507 at $0.30 and $3. Qwen3.8 Flash takes the larger context window at 1M tokens, compared with 256K for Qwen3-235B-A22B-Thinking-2507. Qwen3-235B-A22B-Thinking-2507 accepts text as input, while Qwen3.8 Flash accepts text, images, and video. On tooling, only Qwen3-235B-A22B-Thinking-2507 supports function calling and only Qwen3-235B-A22B-Thinking-2507 offers structured output.

CharacteristicQwen3-235B-A22B-Thinking-2507Qwen3.8 Flash
CompanyAlibabaAlibaba
Release DateJuly 24, 2025August 26, 2026
Parameters235B125B
MultimodalNoYes
Context (input)256K1.0M
Context (output)33K131K
Input Price / 1M$0.30$0.15
Output Price / 1M$3.00$0.47
Average Score69.2%68.5%
Benchmarks
LiveCodeBench v674.1%91.9%
Humanity's Last Exam18.2%35.9%
GPQA81.1%91.7%

Visual Benchmark Comparison

Qwen3-235B-A22B-Thinking-2507
Qwen3.8 Flash
LiveCodeBench v60.7 vs 0.9
0.7
0.9
Humanity's Last Exam0.2 vs 0.4
0.2
0.4
GPQA0.8 vs 0.9
0.8
0.9

Verdict

Qwen3.8 Flash leads in 3 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: Qwen3-235B-A22B-Thinking-2507 — 0.7, Qwen3.8 Flash — 0.7.

API Cost

Qwen3.8 Flash is 5.3x cheaper: input $0.15/1M vs $0.30/1M tokens.

Context Window

Qwen3.8 Flash supports a larger context: 1M vs 256K tokens.

Recency

Qwen3.8 Flash is newer: released 8/26/2026 vs 7/24/2025.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — Qwen3-235B-A22B-Thinking-2507 or Qwen3.8 Flash?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — Qwen3-235B-A22B-Thinking-2507 or Qwen3.8 Flash?
Qwen3.8 Flash is cheaper for input: $0.15 per 1M tokens vs $0.30.
Which has a larger context window — Qwen3-235B-A22B-Thinking-2507 or Qwen3.8 Flash?
Qwen3.8 Flash supports a larger context: 1,048,576 tokens vs 256,000.

The Qwen3-235B-A22B-Thinking-2507 and Qwen3.8 Flash comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the Qwen3-235B-A22B-Thinking-2507 or Qwen3.8 Flash page. See also the complete list of AI model comparisons.