ByteDance logo

Seed 1.8

Multimodal
ByteDance

Seed 1.8 is ByteDance's proprietary multimodal agent model, released 2026-02-17 and optimized for long-horizon agent scenarios with upgraded multimodal comprehension and more flexible context management. It accepts text, image, and video input and reasons before answering. The vendor report lists 94.3% on AIME 2025, 92.3% on MMLU, 84.9% on MMLU-Pro, 83.4% on MMMU, 81.3% on MathVista, 72.9% on SWE-Bench Verified, and 70.7% on AndroidWorld GUI control. DeepInfra serves it with a 256K-token context at $0.25 / $2.00 per 1M input/output tokens.

Key Specifications

Parameters
-
Context
256.0K
Release Date
February 17, 2026
Average Score
68.3%

Timeline

Key dates in the model's history
Announcement
February 17, 2026
Last Update
September 11, 2026
Today
September 19, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.25
Output (per 1M tokens)
$2.00
Max Input Tokens
256.0K
Max Output Tokens
256.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Official Table 1; think-high; no tools; Pass@1Self-reported
92.3%

Programming

Programming skills tests
SWE-Bench Verified
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
72.9%

Multimodal

Working with images and visual data
AI2D
Official Table 2; think-high; Pass@1Self-reported
89.1%
MathVista
Official Table 2; think-high; Pass@1Self-reported
87.7%
MMMU
Official Table 2; think-high; Pass@1Self-reported
83.4%

Other Tests

Specialized benchmarks
AetherCode
Official Table 1; think-high; no tools; Pass@1Self-reported
38.2%
AIME 2025
Official Table 1; think-high; no tools; Pass@1Self-reported
94.3%
AMO Bench
Official Table 1; think-high; no tools; Pass@1Self-reported
60.0%
AndroidWorld
Official Table 5; think-high; agentic GUI evaluationSelf-reported
70.7%
ARC-AGI
Official Table 1; think-high; no tools; Pass@1Self-reported
67.9%
Beyond AIME
Official Table 1; think-high; no tools; Pass@1Self-reported
77.0%
BFCL-V4
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
57.2%
BLINK
Official Table 2; think-high; Pass@1Self-reported
74.3%
BrowseComp
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
67.6%
BrowseComp-zh
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
81.3%
CharXiv-R
Official Table 2; think-high; Pass@1Self-reported
71.4%
CountBench
Official Table 2; think-high; Pass@1Self-reported
96.3%
DUDE
Official Table 2; think-high; Pass@1Self-reported
69.4%
DynaMath
Official Table 2; think-high; Pass@1Self-reported
61.5%
EMMA
Official Table 2; think-high; Pass@1Self-reported
60.9%
ERQA
Official Table 2; think-high; Pass@1Self-reported
58.8%
Hallusion Bench
Official Table 2; think-high; Pass@1Self-reported
63.9%
Humanity's Last Exam (with tools, text-only)
Official Table 4; think-high; agentic search; text-only with toolsSelf-reported
40.9%
LiveCodeBench Pro
Official Table 1; think-high; no tools; Elo ratingSelf-reported
64.3%
LiveCodeBench v6
Official Table 1; think-high; no tools; Pass@1Self-reported
79.5%
LiveSports-3K
Official Table 3; think-high; public video evaluationSelf-reported
77.5%
LongVideoBench
Official Table 3; think-high; public video evaluationSelf-reported
77.4%
LVBench
Official Table 3; think-high; public video evaluationSelf-reported
73.0%
MathVision
Official Table 2; think-high; Pass@1Self-reported
81.3%
Minerva
Official Table 3; think-high; public video evaluationSelf-reported
62.4%
MMLU-Pro
Official Table 1; think-high; no tools; Pass@1Self-reported
84.9%
MMMU-Pro
Official Table 2; think-high; Pass@1Self-reported
73.2%
MMStar
Official Table 2; think-high; Pass@1Self-reported
79.9%
MMVU
Official Table 3; think-high; public video evaluationSelf-reported
73.1%
MotionBench
Official Table 3; think-high; public video evaluationSelf-reported
70.6%
MuirBench
Official Table 2; think-high; Pass@1Self-reported
78.7%
Multi-Challenge
Official Table 1; think-high; no tools; Pass@1Self-reported
66.7%
Multi-SWE-Bench
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
42.0%
OmniDocBench 1.5
Official Table 2; think-high; normalized edit distance (lower is better)Self-reported
10.6%
OSWorld
Official Table 5; think-high; agentic GUI evaluationSelf-reported
61.9%
OVBench
Official Table 3; think-high; public video evaluationSelf-reported
65.1%
OVOBench
Official Table 3; think-high; public video evaluationSelf-reported
72.6%
PHYBench
Official Table 1; think-high; no tools; Pass@1Self-reported
41.0%
RefSpatialBench
Official Table 2; think-high; Pass@1Self-reported
56.3%
SimpleVQA
Official Table 2; think-high; Pass@1Self-reported
65.4%
SuperGPQA
Official Table 1; think-high; no tools; Pass@1Self-reported
64.8%
TempCompass
Official Table 3; think-high; public video evaluationSelf-reported
86.9%
Terminal-Bench 2.0
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
45.2%
TOMATO
Official Table 3; think-high; public video evaluationSelf-reported
60.8%
TVBench
Official Table 3; think-high; public video evaluationSelf-reported
71.5%
VideoMME w sub.
Official Table 3; think-high; subtitles includedSelf-reported
87.8%
VideoMMMU
Official Table 3; think-high; public video evaluationSelf-reported
82.7%
VideoSimpleQA
Official Table 3; think-high; public video evaluationSelf-reported
67.8%
VLMsAreBiased
Official Table 2; think-high; Pass@1Self-reported
62.0%
VLMsAreBlind
Official Table 2; think-high; Pass@1Self-reported
93.0%
WideSearch
Official Table 4; think-high; benchmark-specific agent toolsSelf-reported
63.8%
ZEROBench
Official Table 2; think-high; Pass@1Self-reported
11.0%

License & Metadata

License
proprietary
Announcement Date
February 17, 2026
Last Updated
September 11, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.