StepFun logo

Step 3.7 Flash

Multimodal
StepFun

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro. It is served by deepinfra with a 256K-token context window at $0.2 / $1.15 per 1M input/output tokens.

Key Specifications

Parameters
198.0B
Context
262.1K
Release Date
June 10, 2026
Average Score
52.2%

Timeline

Key dates in the model's history
Announcement
June 10, 2026
Last Update
September 12, 2026
Today
September 22, 2026

Technical Specifications

Parameters
198.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.20
Output (per 1M tokens)
$1.15
Max Input Tokens
262.1K
Max Output Tokens
262.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
GDPval
StepFun internal pairwise evaluation across GDPval's 44 occupations; no further sampling details are disclosed.Self-reported
45.8%
Humanity's Last Exam (with tools, text-only)
Humanity's Last Exam with tool use; the page classifies this among non-multimodal evaluations and provides no tool-set, harness, or sampling details.Self-reported
47.2%
SWE-Bench Pro
Publisher-reported Agentic Coding comparison result; no harness, sampling count, or pass@k details are disclosed.Self-reported
56.3%
Terminal-Bench 2.1
StepFun internal evaluation on Terminal-Bench 2.1; no harness or sampling details are disclosed.Self-reported
59.5%

License & Metadata

License
apache_2_0
Announcement Date
June 10, 2026
Last Updated
September 12, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.