Alibaba logo

Qwen3 VL 4B Thinking

Multimodal
Alibaba

Qwen3 VL 4B Thinking is the reasoning variant of the compact Qwen3 VL 4B vision-language model, released in September 2025. It applies extended chain-of-thought to multimodal understanding tasks, scoring 0.94 on DocVQA and 0.87 on MMBench-V1.1. It is Apache 2.0 licensed and open-weight.

Key Specifications

Parameters
4.0B
Context
262.1K
Release Date
September 22, 2025
Average Score
65.8%

Timeline

Key dates in the model's history
Announcement
September 22, 2025
Last Update
August 27, 2026
Today
September 10, 2026

Technical Specifications

Parameters
4.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$1.00
Max Input Tokens
262.1K
Max Output Tokens
262.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Self-reported
81.5%

Reasoning

Logical reasoning and analysis
GPQA
Self-reported
64.1%

Multimodal

Working with images and visual data
AI2D
TESTSelf-reported
84.9%

Other Tests

Specialized benchmarks
DocVQA
Self-reported by the model providerSelf-reported
94.0%
ScreenSpot
Self-reported by the model providerSelf-reported
93.0%
MMBench-V1.1
en, V1.1Self-reported
87.0%
AIME 2025
Self-reported
74.5%
Arena-Hard v2
winrateSelf-reported
36.8%
BFCL-v3
Self-reported
67.3%
BLINK
Self-reported
63.4%
CC-OCR
overallSelf-reported
73.8%
CharadesSTA
Self-reported
59.0%
CharXiv-D
DQSelf-reported
83.9%
CharXiv-R
RQSelf-reported
50.3%
Creative Writing v3
Self-reported
76.1%
ERQA
Self-reported
47.3%
Hallusion Bench
Self-reported
64.1%
HMMT25
Self-reported
53.1%
IFEval
Self-reported
82.6%
Include
Self-reported
64.6%
InfoVQAtest
Self-reported
83.0%
LiveBench 20241125
Self-reported
68.4%
LiveCodeBench v6
25.02-25.05Self-reported
51.3%
LVBench
Self-reported
53.5%
MathVision
Self-reported
60.0%
MathVista-Mini
Self-reported
79.5%
MLVU-M
MCQSelf-reported
75.7%
MM-MT-Bench
Self-reported
7.7%
MMLU-Pro
Self-reported
73.6%
MMLU-ProX
Self-reported
65.0%
MMLU-Redux
Self-reported
86.0%
MMMU (val)
AISelf-reported
70.8%
MMMU-Pro
fullSelf-reported
57.0%
MMStar
Self-reported
73.2%
MuirBench
Self-reported
75.0%
Multi-IF
Self-reported
73.6%
MVBench
Self-reported
69.3%
OCRBench
Self-reported
80.8%
OCRBench-V2 (en)
en/zh - using en valueSelf-reported
61.8%
OCRBench-V2 (zh)
zhSelf-reported
55.8%
ODinW
13Self-reported
39.4%
OSWorld
Self-reported
31.4%
PolyMATH
Self-reported
44.6%
RealWorldQA
Self-reported
73.2%
ScreenSpot Pro
Self-reported
49.2%
SuperGPQA
Self-reported
46.8%
VideoMMMU
Self-reported
69.4%
WritingBench
Self-reported
84.0%

License & Metadata

License
apache-2.0
Announcement Date
September 22, 2025
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.