Step3-VL-10B
MultimodalSTEP3-VL-10B is a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Built on a unified, fully unfrozen pre-training strategy on 1.2T multimodal tokens integrating a language-aligned Perception Encoder with a Qwen3-8B decoder. Features Parallel Coordinated Reasoning (PaCoRe) to scale test-time compute for complex perceptual reasoning.
Key Specifications
Timeline
Technical Specifications
Benchmark Results
Model performance metrics across various tests and benchmarks
Multimodal
Other Tests
License & Metadata
Similar Models
All ModelsStep-3.5-Flash
StepFun
Step 3.7 Flash
StepFun
EXAONE 4.5 33B
LG AI Research
Qwen3.6-27B
Alibaba
Qwen2-VL-72B-Instruct
Alibaba
Gemma 4 31B
Qwen3 VL 32B Thinking
Alibaba
Llama 3.2 90B Instruct
Meta
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.