Step-3.5-Flash
MultimodalStep 3.5 Flash is a fast and efficient language model by StepFun, optimized for low-latency inference while maintaining strong performance in reasoning, coding, and general knowledge tasks. It offers an excellent balance between speed and quality for production deployments.
Key Specifications
Parameters
196.0B
Context
65.5K
Release Date
February 1, 2026
Average Score
77.6%
Timeline
Key dates in the model's history
Announcement
February 1, 2026
Last Update
February 12, 2026
Today
September 4, 2026
Technical Specifications
Parameters
196.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.40
Max Input Tokens
65.5K
Max Output Tokens
8.2K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
• Self-reported
Other Tests
Specialized benchmarks
AIME 2025
• Self-reported
Tau-bench
• Self-reported
LiveCodeBench v6
• Self-reported
IMO-AnswerBench
• Self-reported
BrowseComp
Context Manager • Self-reported
Terminal-Bench 2.0
• Self-reported
License & Metadata
License
apache-2.0
Announcement Date
February 1, 2026
Last Updated
February 12, 2026
Similar Models
All ModelsInkling-Small
Thinking Machines Lab
MM276.0B
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$0.30/1M tokens
GPT OSS 120B
OpenAI
MM120.0B
Best score:0.9 (MMLU)
Released:Aug 2025
Price:$0.15/1M tokens
Command A+
Cohere
MM218.0B
Best score:0.8 (TAU)
Released:May 2026
GLM-4.6
Zhipu AI
MM357.0B
Best score:0.8 (GPQA)
Released:Sep 2025
Price:$0.60/1M tokens
Qwen3.8 Flash
Alibaba
MM125.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.15/1M tokens
Kimi K3
Moonshot AI
MM2.8T
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$3.00/1M tokens
Qwen3.8 Flash Next
Alibaba
MM180.0B
Best score:0.9 (GPQA)
Released:Aug 2026
Price:$0.16/1M tokens
Kimi K2.5
Moonshot AI
MM1.0T
Best score:0.9 (GPQA)
Released:Jan 2026
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.