Alibaba logo

Qwen3 VL 4B Instruct

Multimodal
Alibaba

Qwen3 VL 4B Instruct is a compact 4B-parameter multimodal vision-language model from the Qwen team, released in September 2025. It provides strong document and GUI understanding for its size, scoring 0.95 on DocVQA and 0.94 on ScreenSpot. It is Apache 2.0 licensed and open-weight.

Key Specifications

Parameters
4.0B
Context
262.1K
Release Date
September 22, 2025
Average Score
61.7%

Timeline

Key dates in the model's history
Announcement
September 22, 2025
Last Update
August 27, 2026
Today
September 10, 2026

Technical Specifications

Parameters
4.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.60
Max Input Tokens
262.1K
Max Output Tokens
262.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Self-reported
77.2%

Multimodal

Working with images and visual data
AI2D
TESTSelf-reported
84.1%

Other Tests

Specialized benchmarks
DocVQA
Self-reported by the model providerSelf-reported
95.0%
ScreenSpot
Self-reported by the model providerSelf-reported
94.0%
OCRBench
Self-reported by the model providerSelf-reported
88.0%
AIME 2025
Self-reported
46.6%
BFCL-v3
Self-reported
63.3%
BLINK
Self-reported
65.8%
CC-OCR
overallSelf-reported
76.2%
CharadesSTA
Self-reported
55.5%
CharXiv-D
DQSelf-reported
76.2%
CharXiv-R
RQSelf-reported
39.7%
ERQA
Self-reported
41.3%
Hallusion Bench
Self-reported
57.6%
HMMT25
Self-reported
30.7%
IFEval
Self-reported
82.3%
Include
Self-reported
61.4%
InfoVQAtest
Self-reported
80.3%
LiveBench 20241125
Self-reported
60.9%
LiveCodeBench v6
25.02-25.05Self-reported
37.9%
LVBench
Self-reported
56.2%
MathVision
Self-reported
51.6%
MathVista-Mini
Self-reported
73.7%
MLVU-M
MCQSelf-reported
75.3%
MM-MT-Bench
Self-reported
7.5%
MMBench-V1.1
en, V1.1Self-reported
85.1%
MMLU-Pro
Self-reported
67.1%
MMLU-ProX
Self-reported
59.4%
MMLU-Redux
Self-reported
81.5%
MMMU (val)
AISelf-reported
67.4%
MMMU-Pro
fullSelf-reported
53.2%
MMStar
Self-reported
69.8%
MuirBench
Self-reported
63.8%
MVBench
Self-reported
68.9%
OCRBench-V2 (en)
en/zh - using en valueSelf-reported
63.7%
OCRBench-V2 (zh)
zhSelf-reported
57.6%
ODinW
13Self-reported
48.2%
OSWorld
Self-reported
26.2%
PolyMATH
Self-reported
28.8%
RealWorldQA
Self-reported
70.9%
ScreenSpot Pro
Self-reported
59.5%
SimpleQA
VQASelf-reported
48.0%
SuperGPQA
Self-reported
40.3%
VideoMMMU
Self-reported
56.2%
WritingBench
Self-reported
82.5%

License & Metadata

License
apache-2.0
Announcement Date
September 22, 2025
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.