StepFun logo

Step3-VL-10B

Multimodal
StepFun

STEP3-VL-10B is a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Built on a unified, fully unfrozen pre-training strategy on 1.2T multimodal tokens integrating a language-aligned Perception Encoder with a Qwen3-8B decoder. Features Parallel Coordinated Reasoning (PaCoRe) to scale test-time compute for complex perceptual reasoning.

Key Specifications

Parameters
10.0B
Context
-
Release Date
January 15, 2026
Average Score
79.2%

Timeline

Key dates in the model's history
Announcement
January 15, 2026
Last Update
September 10, 2026
Today
September 22, 2026

Technical Specifications

Parameters
10.0B
Training Tokens
1.2T tokens
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Benchmark Results

Model performance metrics across various tests and benchmarks

Multimodal

Working with images and visual data
MathVista
Self-reported
84.0%
MMMU
Self-reported
78.1%

Other Tests

Specialized benchmarks
AIME 2025
Self-reported
87.7%
MathVision
Self-reported
70.8%
MMBench
Self-reported
91.8%
Multi-Challenge
Self-reported
62.6%

License & Metadata

License
apache_2_0
Announcement Date
January 15, 2026
Last Updated
September 10, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.