Alibaba logo

Qwen3-235B-A22B-Thinking-2507

Alibaba

Qwen3-235B-A22B-Thinking-2507 is an advanced thinking-mode model on the Mixture-of-Experts architecture with 235B total parameters (22B active). Features 94 layers, 128 experts (8 active), and supports a native context length of 262K. This version is significantly improved in reasoning, achieving leading results among open models with thinking mode in logic, math, science, coding, and academic benchmarks.

Key Specifications

Parameters
235.0B
Context
256.0K
Release Date
July 24, 2025
Average Score
69.2%

Timeline

Key dates in the model's history
Announcement
July 24, 2025
Last Update
January 22, 2026
Today
September 10, 2026

Technical Specifications

Parameters
235.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.30
Output (per 1M tokens)
$3.00
Max Input Tokens
256.0K
Max Output Tokens
131.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
Self-reported by the model providerSelf-reported
81.1%

Other Tests

Specialized benchmarks
MMLU-Redux
Self-reported
94.0%
AIME 2025
Self-reported
92.0%
IFEval
Self-reported
88.0%
Arena-Hard v2
GPT-4 evaluated win ratesSelf-reported
79.7%
BFCL-v3
Self-reported by the model providerSelf-reported
71.9%
CFEval
Self-reported by the model providerSelf-reported
21.3%
Creative Writing v3
Self-reported by the model providerSelf-reported
86.1%
HMMT25
Self-reported by the model providerSelf-reported
83.9%
Humanity's Last Exam
text-only subsetSelf-reported
18.2%
Include
Self-reported by the model providerSelf-reported
81.0%
LiveBench 20241125
Self-reported by the model providerSelf-reported
78.4%
LiveCodeBench v6
25.02-25.05Self-reported
74.1%
MMLU-Pro
Self-reported by the model providerSelf-reported
84.4%
MMLU-ProX
Self-reported by the model providerSelf-reported
81.0%
Multi-IF
Self-reported by the model providerSelf-reported
80.6%
OJBench
Self-reported by the model providerSelf-reported
32.5%
PolyMATH
Self-reported by the model providerSelf-reported
60.1%
SuperGPQA
Self-reported by the model providerSelf-reported
64.9%
Tau2 Airline
Self-reported by the model providerSelf-reported
58.0%
Tau2 Retail
Self-reported by the model providerSelf-reported
71.9%
Tau2 Telecom
Self-reported by the model providerSelf-reported
45.6%
TAU-bench Airline
Self-reported by the model providerSelf-reported
46.0%
TAU-bench Retail
Self-reported by the model providerSelf-reported
67.8%
WritingBench
Self-reported by the model providerSelf-reported
88.3%

License & Metadata

License
apache-2.0
Announcement Date
July 24, 2025
Last Updated
January 22, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.