GPT-5.5
MultimodalGPT-5.5 is OpenAI's smartest model yet, designed for real work across agentic coding, computer use, knowledge work, and early scientific research. It matches GPT-5.4 per-token latency in real-world serving while reaching a much higher level of intelligence and using significantly fewer tokens to complete the same tasks, with a 1M-token context window in the API.
Key Specifications
Parameters
-
Context
1.1M
Release Date
April 23, 2026
Average Score
66.4%
Timeline
Key dates in the model's history
Announcement
April 23, 2026
Today
August 27, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
December 1, 2025
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$30.00
Max Input Tokens
1.1M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Reasoning
Logical reasoning and analysis
GPQA
GPQA Diamond. Reasoning effort xhigh. • Self-reported
Other Tests
Specialized benchmarks
ARC-AGI
ARC-AGI-1 (Verified). Reasoning effort xhigh. • Self-reported
ARC-AGI v2
ARC-AGI-2 (Verified). Reasoning effort xhigh. • Self-reported
BixBench
BixBench bioinformatics data analysis. Reasoning effort xhigh. • Self-reported
BrowseComp
BrowseComp agentic web browsing benchmark. Reasoning effort xhigh. • Self-reported
CyberGym
CyberGym. Reasoning effort xhigh. • Self-reported
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, xhigh effort; Pass@1 67% ± 6%. • Self-reported
Finance Agent
FinanceAgent v1.1. Reasoning effort xhigh. • Self-reported
Finance Agent v2
• Self-reported
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at xhigh effort: 43.0%. • Self-reported
FrontierMath
FrontierMath Tier 4. Reasoning effort xhigh. • Self-reported
FrontierSWE
Codex • Self-reported
GDPval-MM
GDPval (wins or ties). Reasoning effort xhigh. • Self-reported
GeneBench
GeneBench. Reasoning effort xhigh. • Self-reported
Graphwalks BFS >128k
Graphwalks BFS 1M f1. Reasoning effort xhigh. • Self-reported
Graphwalks parents >128k
Graphwalks parents 1M f1. Reasoning effort xhigh. • Self-reported
Humanity's Last Exam
Humanity's Last Exam (with tools). Reasoning effort xhigh. • Self-reported
Legal Agent Benchmark
All-pass rate (Harvey held-out set) • Self-reported
LiveBench
2026-01-08, Thinking xHigh Effort • Self-reported
MCP Atlas
MCP Atlas (Scale AI April 2026 update). Reasoning effort xhigh. • Self-reported
MMMU-Pro
MMMU Pro (with tools). Reasoning effort xhigh. • Self-reported
MRCR v2 (8-needle)
OpenAI MRCR v2 8-needle, 512K-1M context range. Reasoning effort xhigh. • Self-reported
OfficeQA Pro
OfficeQA Pro. Reasoning effort xhigh. • Self-reported
OSWorld-Verified
OSWorld-Verified desktop navigation benchmark. Reasoning effort xhigh. • Self-reported
SWE-Bench Pro
SWE-Bench Pro (Public). Reasoning effort xhigh. • Self-reported
Tau2 Telecom
Tau2-bench Telecom (original prompts, no prompt tuning). Reasoning effort xhigh. • Self-reported
Terminal-Bench 2.0
Terminal-Bench 2.0. Reasoning effort xhigh. • Self-reported
Toolathlon
Toolathlon tool-use agentic tasks benchmark. Reasoning effort xhigh. • Self-reported
License & Metadata
License
proprietary
Announcement Date
April 23, 2026
Last Updated
August 27, 2026
Similar Models
All ModelsGPT-5.2
OpenAI
MM
Best score:1.0 (TAU)
Released:Dec 2025
Price:$1.50/1M tokens
GPT-5.4
OpenAI
MM
Best score:1.0 (TAU)
Released:Mar 2026
Price:$2.50/1M tokens
GPT-5 Medium
OpenAI
MM
Best score:0.9 (GPQA)
Released:Aug 2025
Price:$0.75/1M tokens
GPT-5.1 Instant
OpenAI
MM
Best score:1.0 (TAU)
Released:Nov 2025
Price:$0.30/1M tokens
GPT-5.4 mini
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.42/1M tokens
GPT-5.4 nano
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.12/1M tokens
GPT-5.6 Sol
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$5.00/1M tokens
GPT-5.6 Terra
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$2.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.