Alibaba logo

Qwen3 VL 4B Instruct

Multimodal
Alibaba

Qwen3 VL 4B Instruct is a compact 4B-parameter multimodal vision-language model from the Qwen team, released in September 2025. It provides strong document and GUI understanding for its size, scoring 0.95 on DocVQA and 0.94 on ScreenSpot. It is Apache 2.0 licensed and open-weight.

Key Specifications

Parameters
4.0B
Context
262.1K
Release Date
September 22, 2025
Average Score
92.3%

Timeline

Key dates in the model's history
Announcement
September 22, 2025
Last Update
August 27, 2026

Technical Specifications

Parameters
4.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.60
Max Input Tokens
262.1K
Max Output Tokens
262.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
DocVQA
Self-reported by the model providerSelf-reported
95.0%
ScreenSpot
Self-reported by the model providerSelf-reported
94.0%
OCRBench
Self-reported by the model providerSelf-reported
88.0%

License & Metadata

License
apache-2.0
Announcement Date
September 22, 2025
Last Updated
August 27, 2026

Compare Qwen3 VL 4B Instruct

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.