Key Specifications
Parameters
-
Context
256.0K
Release Date
February 17, 2026
Average Score
58.5%
Timeline
Key dates in the model's history
Announcement
February 17, 2026
Last Update
September 10, 2026
Today
September 19, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.40
Max Input Tokens
256.0K
Max Output Tokens
256.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Reasoning
Logical reasoning and analysis
GPQA
Model Card Table 4 efficient-models; GPQA Diamond percentage score mapped to existing GPQA convention. • Self-reported
Multimodal
Working with images and visual data
MMMU
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
Other Tests
Specialized benchmarks
AetherCode
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
AIME 2025
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
AIME 2026
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
ARC-AGI
Model Card Table 8; ArcAGI1-Image percentage score. • Self-reported
ArcAGI2
Model Card Table 4 efficient-models; ARC-AGI-2 text/general-reasoning percentage score. • Self-reported
ARC-AGI v2
Model Card Table 8; ArcAGI2-Image percentage score. • Self-reported
BabyVision
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
Beyond AIME
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
BLINK
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
ChartQAPro
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
CharXiv-D
Model Card Table 8; CharXiv-DQ percentage score. • Self-reported
CharXiv-R
Model Card Table 8; CharXiv-RQ percentage score. • Self-reported
CL-bench
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
CodeForces
Model Card Table 4 efficient-models; Codeforces Elo rating. • Self-reported
COLLIE
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
ContPhy
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
CountBench
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
CrossVid
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
DUDE
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
DynaMath
Model Card Table 8; DynaMath worst-case accuracy over 10 variants. • Self-reported
EMMA
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
ERQA
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
FActScore
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
FRAMES
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
FrontierScience Olympiad
Model Card Table 4 efficient-models; FrontierScience olympiad split percentage score. • Self-reported
FrontierScience Research
Model Card Table 4 efficient-models; FrontierScience research split percentage score. • Self-reported
Global PIQA
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
Graphwalks BFS <128k
Model Card Table 4 efficient-models; Graphwalks BFS <128K percentage score. • Self-reported
Graphwalks parents <128k
Model Card Table 4 efficient-models; Graphwalks Parents <128K percentage score. • Self-reported
Hallusion Bench
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
HealthBench
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
HealthBench Hard
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
HiPhO
Model Card Table 8; HiPhO average normalized score over 13 physics olympiad competitions. • Self-reported
Humanity's Last Exam (no tools, text-only)
Model Card Table 4 efficient-models; no-tools, text-only HLE percentage score. • Self-reported
LiveCodeBench v6
Model Card Table 4 efficient-models; LiveCodeBench v6 percentage score. • Self-reported
LiveSports-3K
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
LongFact Concepts
Model Card Table 4 efficient-models; LongFact Concepts percentage score. • Self-reported
LongFact Objects
Model Card Table 4 efficient-models; LongFact Objects percentage score. • Self-reported
LongVideoBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
LVBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
MathArena Apex
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
MathVision
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
MathVista-Mini
Model Card Table 8; MathVista testmini Pass@1 percentage score. • Self-reported
Minerva
Model Card Table 9 public video benchmarks; subtitle-marked Minerva score. • Self-reported
MMLongBench-Doc
Model Card Table 8; MMLongBench-Doc percentage score. • Self-reported
MMLU-Pro
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
MMMLU
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
MMMU-Pro
Model Card Table 8; MMMU-Pro aggregate over Standard 10-option and Vision subsets. • Self-reported
MMStar
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
MMVU
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
MotionBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
MRCR v2 (8-needle)
Model Card Table 4 efficient-models; MRCR v2 8-needle percentage score. • Self-reported
MTVQA
Model Card Table 8; MTVQA scored by DeepSeek-V3-0324 LLM judge. • Self-reported
MuirBench
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
Multi-Challenge
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
OCRBench_V2
Model Card Table 8; OCRBenchv2 average of overall English and Chinese scores. • Self-reported
OmniDocBench 1.5
Model Card Table 8; OmniDocBench 1.5 overall edit distance. • Self-reported
OVBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
OVOBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
PHYBench
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
RealWorldQA
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
RefSpatialBench
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
SimpleQA Verified
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
SimpleVQA
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
SuperChem
Model Card Table 4 efficient-models; SuperChem text-only percentage score. • Self-reported
SuperGPQA
Model Card Table 4 efficient-models; reported percentage score. • Self-reported
TempCompass
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
TOMATO
Model Card Table 9 public video benchmarks; default TOMATO score. • Self-reported
TreeBench
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
TVBench
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
VideoHolmes
Model Card Table 9 public video benchmarks; subtitle-marked VideoHolmes score. • Self-reported
VideoMME w sub.
Model Card Table 9 public video benchmarks; VideoMME with subtitles. • Self-reported
VideoMMMU
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
VideoSimpleQA
Model Card Table 9 public video benchmarks; reported percentage score. • Self-reported
VisFactor
Model Card Table 8; VisFactor macro-average percentage score. • Self-reported
VisuLogic
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
VLMsAreBiased
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
VLMsAreBlind
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
WorldVQA
Model Card Table 8 public visual-language benchmarks; Pass@1 percentage score. • Self-reported
ZEROBench
Model Card Table 8; ZEROBench main split percentage score. • Self-reported
ZEROBench-Sub
Model Card Table 8; ZEROBench sub split percentage score. • Self-reported
License & Metadata
License
proprietary
Announcement Date
February 17, 2026
Last Updated
September 10, 2026
Similar Models
All ModelsSeed 2.0 Lite
ByteDance
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Seed 2.1 Pro
ByteDance
MM
Released:Jun 2026
Seed 2.1 Turbo
ByteDance
MM
Released:Jun 2026
Seed 1.8
ByteDance
MM
Best score:0.9 (MMLU)
Released:Feb 2026
Price:$0.25/1M tokens
Seedance 1.5 Pro
ByteDance
MM
Released:Jan 1970
Seed 2.0 Pro
ByteDance
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$0.50/1M tokens
ERNIE 5.0
Baidu
MM
Best score:0.8 (GPQA)
Released:Jan 2025
Grok-4
xAI
MM
Best score:0.9 (GPQA)
Released:Jul 2025
Price:$3.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.