Alibaba logo

Qwen3.8 Flash Next

Multimodal
Alibaba

Qwen3.8 Flash Next is an open-weight multimodal MoE from Alibaba's Qwen team, released August 2026 as an early preview of the Qwen4 architecture. A 125B-parameter backbone with 51B N-gram embeddings and a 4B multi-token prediction module activates only 6B parameters per token. It introduces a GDN + Qwen Sparse Attention hybrid, Gated Residual and N-gram Embedding, and reportedly costs about one-ninth of Qwen3.7-Plus to train. Native 262K context, extensible to 1M with YaRN. Licensed under the Qwen community license (not Apache-2.0).

Key Specifications

Parameters
180.0B
Context
1.0M
Release Date
August 26, 2026
Average Score
67.1%

Timeline

Key dates in the model's history
Announcement
August 26, 2026
Last Update
August 27, 2026
Today
September 8, 2026

Technical Specifications

Parameters
180.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.16
Output (per 1M tokens)
$0.47
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA Diamond
GPQA Diamond, scientific reasoningSelf-reported
91.7%

Other Tests

Specialized benchmarks
DeepSWE 1.1
Agentic coding, Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, 256K contextSelf-reported
58.7%
SWE-bench Pro
Claude Code harness, temp=1.0, top_p=0.95, 256K contextSelf-reported
62.5%
SWE-bench Multilingual
mini-SWE-agent harness, temp=1.0, top_p=0.95, 256K contextSelf-reported
81.0%
LiveCodeBench v6
Competitive coding, LiveCodeBench v6Self-reported
91.9%
Agents' Last Exam
Frontier agentic tasks; Pass@1: 24.3, Score: 51.2Self-reported
51.2%
Toolathlon Verified
Real-world tool use, Pass@1Self-reported
73.5%
IFBench
Instruction followingSelf-reported
81.3%
Humanity's Last Exam
Multidisciplinary reasoning, judged by GPT-4oSelf-reported
35.9%
AndroidWorld
Self-reported by the model providerSelf-reported
84.5%
CharXiv-R
With code interpreterSelf-reported
90.6%
ClawEval-MM
Average score across three trialsSelf-reported
60.4%
CoWorkBench
Self-reported by the model providerSelf-reported
73.9%
ERQA
Self-reported by the model providerSelf-reported
72.3%
Humanity's Last Exam
Judged by GPT-4oSelf-reported
35.9%
Job Bench
Self-reported by the model providerSelf-reported
55.7%
LVBench
Self-reported by the model providerSelf-reported
76.6%
MathVision
With code interpreterSelf-reported
95.7%
NL2Repo
Claude Code harnessSelf-reported
48.1%
OSWorld 2.0
Binary completionSelf-reported
19.4%
RealWorldQA
Self-reported by the model providerSelf-reported
88.5%
RecreationBench
Self-reported by the model providerSelf-reported
49.9%
Vision2Web
Claude Code harness, judged by gpt-5.4-2026-03-05Self-reported
64.0%

License & Metadata

License
qwen
Announcement Date
August 26, 2026
Last Updated
August 27, 2026

Compare Qwen3.8 Flash Next

All comparisons

Articles about Qwen3.8 Flash Next

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.