Qwen3 VL 4B Thinking
MultimodalQwen3 VL 4B Thinking is the reasoning variant of the compact Qwen3 VL 4B vision-language model, released in September 2025. It applies extended chain-of-thought to multimodal understanding tasks, scoring 0.94 on DocVQA and 0.87 on MMBench-V1.1. It is Apache 2.0 licensed and open-weight.
Key Specifications
Parameters
4.0B
Context
262.1K
Release Date
September 22, 2025
Average Score
65.8%
Timeline
Key dates in the model's history
Announcement
September 22, 2025
Last Update
August 27, 2026
Today
September 10, 2026
Technical Specifications
Parameters
4.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$1.00
Max Input Tokens
262.1K
Max Output Tokens
262.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
• Self-reported
Reasoning
Logical reasoning and analysis
GPQA
• Self-reported
Multimodal
Working with images and visual data
AI2D
TEST • Self-reported
Other Tests
Specialized benchmarks
DocVQA
Self-reported by the model provider • Self-reported
ScreenSpot
Self-reported by the model provider • Self-reported
MMBench-V1.1
en, V1.1 • Self-reported
AIME 2025
• Self-reported
Arena-Hard v2
winrate • Self-reported
BFCL-v3
• Self-reported
BLINK
• Self-reported
CC-OCR
overall • Self-reported
CharadesSTA
• Self-reported
CharXiv-D
DQ • Self-reported
CharXiv-R
RQ • Self-reported
Creative Writing v3
• Self-reported
ERQA
• Self-reported
Hallusion Bench
• Self-reported
HMMT25
• Self-reported
IFEval
• Self-reported
Include
• Self-reported
InfoVQAtest
• Self-reported
LiveBench 20241125
• Self-reported
LiveCodeBench v6
25.02-25.05 • Self-reported
LVBench
• Self-reported
MathVision
• Self-reported
MathVista-Mini
• Self-reported
MLVU-M
MCQ • Self-reported
MM-MT-Bench
• Self-reported
MMLU-Pro
• Self-reported
MMLU-ProX
• Self-reported
MMLU-Redux
• Self-reported
MMMU (val)
AI • Self-reported
MMMU-Pro
full • Self-reported
MMStar
• Self-reported
MuirBench
• Self-reported
Multi-IF
• Self-reported
MVBench
• Self-reported
OCRBench
• Self-reported
OCRBench-V2 (en)
en/zh - using en value • Self-reported
OCRBench-V2 (zh)
zh • Self-reported
ODinW
13 • Self-reported
OSWorld
• Self-reported
PolyMATH
• Self-reported
RealWorldQA
• Self-reported
ScreenSpot Pro
• Self-reported
SuperGPQA
• Self-reported
VideoMMMU
• Self-reported
WritingBench
• Self-reported
License & Metadata
License
apache-2.0
Announcement Date
September 22, 2025
Last Updated
August 27, 2026
Similar Models
All ModelsQwen3.5-4B
Alibaba
MM4.0B
Best score:0.8 (GPQA)
Released:Mar 2026
Qwen3 VL 4B Instruct
Alibaba
MM4.0B
Best score:0.8 (MMLU)
Released:Sep 2025
Price:$0.10/1M tokens
Qwen2.5-Omni-7B
Alibaba
MM7.0B
Best score:0.8 (HumanEval)
Released:Mar 2025
Qwen2.5 VL 7B Instruct
Alibaba
MM8.3B
Released:Jan 2025
Qwen2.5 VL 32B Instruct
Alibaba
MM33.5B
Best score:0.9 (HumanEval)
Released:Feb 2025
Qwen2.5-Coder 7B Instruct
Alibaba
7.0B
Best score:0.9 (HumanEval)
Released:Sep 2024
Granite 3.3 8B Instruct
IBM
MM8.0B
Best score:0.9 (HumanEval)
Released:Apr 2025
Granite 3.3 8B Base
IBM
MM8.2B
Best score:0.9 (HumanEval)
Released:Apr 2025
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.