OpenAI logo

GPT-5.6 Sol

Multimodal
OpenAI

GPT-5.6 Sol is the frontier model in OpenAI's GPT-5.6 family, designed for complex professional work across coding, knowledge work, cybersecurity, and science. It sets state-of-the-art results while using fewer tokens at lower estimated cost, supports max reasoning effort and an ultra multi-agent mode, and has a 1.05M-token context window; the gpt-5.6 alias routes to GPT-5.6 Sol.

Key Specifications

Parameters
-
Context
1.1M
Release Date
July 9, 2026
Average Score
64.5%

Timeline

Key dates in the model's history
Announcement
July 9, 2026
Today
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
February 16, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$30.00
Max Input Tokens
1.1M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond. Max reasoning effort.Self-reported
94.6%

Other Tests

Specialized benchmarks
Agents' Last Exam
Agents' Last Exam. Max reasoning effort.Self-reported
52.7%
ARC-AGI-3
ARC-AGI-3. Max reasoning effort.Self-reported
7.8%
Artificial Analysis
Artificial Analysis Intelligence Index v4.1 (max). Score 59.Self-reported
59.0%
AutomationBench
AutomationBench. Max reasoning effort.Self-reported
18.1%
BenchCAD
BenchCAD (no tools). Max reasoning effort.Self-reported
70.6%
BenchCAD (with Python tool)
BenchCAD with Python tool. Max reasoning effort.Self-reported
83.4%
Big Finance Bench
Big Finance Bench. Max reasoning effort.Self-reported
53.0%
BrowseComp
BrowseComp. Max reasoning effort (single-agent).Self-reported
90.4%
Capture-the-Flag Challenges (Internal)
Capture-the-Flag Challenges (Internal). Max reasoning effort.Self-reported
96.7%
Connectors
Connectors production benchmark pass rate.Self-reported
100.0%
DeepSWE
DeepSWE v1.1. Max reasoning effort.Self-reported
72.7%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 73% ± 3%.Self-reported
73.0%
ExploitBench
ExploitBench. Max reasoning effort.Self-reported
73.5%
ExploitGym
ExploitGym, six-hour cap. Max reasoning effort.Self-reported
33.7%
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at max effort: 47.5%.Self-reported
47.5%
FrontierMath
FrontierMath Tier 1-3 (v2). Max reasoning effort.Self-reported
89.0%
FrontierMath Tier 4 (v2)
FrontierMath Tier 4 (v2). Max reasoning effort.Self-reported
83.0%
GDP.pdf
gdp.pdf. Max reasoning effort.Self-reported
30.7%
GeneBench-Pro
GeneBench Pro. Max reasoning effort.Self-reported
28.7%
Graphwalks BFS >128k
GraphWalks BFS 256k f1. Max reasoning effort.Self-reported
90.7%
Graphwalks BFS 1M
GraphWalks BFS 1M f1. Max reasoning effort.Self-reported
77.1%
HealthBench
HealthBench, length-adjusted score (unadjusted 55.6).Self-reported
57.0%
HealthBench Consensus
HealthBench Consensus, length-adjusted score (unadjusted 95.3).Self-reported
95.5%
HealthBench Hard
HealthBench Hard, length-adjusted score (unadjusted 31.1).Self-reported
33.1%
HealthBench Professional
HealthBench Professional, length-adjusted score (unadjusted 64.1).Self-reported
60.5%
Internal Research Debugging Evaluation
Internal Research Debugging Evaluation. Max reasoning effort.Self-reported
68.3%
KernelGen 1P
KernelGen 1P. Max reasoning effort.Self-reported
61.1%
LifeSciBench
LifeSciBench. Max reasoning effort.Self-reported
59.9%
Management Consulting Tasks (Internal)
Management Consulting Tasks (Internal). Max reasoning effort.Self-reported
43.2%
MedChemBench (Internal)
MedChemBench (Internal). Max reasoning effort.Self-reported
48.3%
MMMU-Pro
MMMU Pro (no tools). Max reasoning effort.Self-reported
83.0%
MMMU-Pro (with tools)
MMMU Pro (with tools). Max reasoning effort.Self-reported
84.6%
MRCR v2 (8-needle)
OpenAI MRCR v2 8-needle, 256K-512K. Max reasoning effort.Self-reported
91.5%
MRCR v2 (8-needle, 512K-1M)
OpenAI MRCR v2 8-needle, 512K-1M. Max reasoning effort.Self-reported
73.8%
NanoGPT
NanoGPT. Max reasoning effort.Self-reported
9.7%
OSWorld 2.0
OSWorld 2.0 binary completion. Max reasoning effort.Self-reported
62.6%
PostTrainBench Lite
PostTrainBench Lite mean reward. Max reasoning effort.Self-reported
50.3%
RSI Index
RSI Index (aggregate self-improvement). Max reasoning effort.Self-reported
57.9%
Search and Function-Calling
Search and Function-Calling production benchmark pass rate.Self-reported
91.0%
SEC-bench Pro
SEC-Bench Pro. Max reasoning effort (single-agent).Self-reported
71.2%
SWE-Bench Pro
SWE-Bench Pro. Max reasoning effort.Self-reported
64.6%
Terminal-Bench 2.1
Terminal-Bench 2.1. Max reasoning effort (single-agent).Self-reported
88.8%
Toolathlon
Toolathlon. Max reasoning effort.Self-reported
58.0%

License & Metadata

License
proprietary
Announcement Date
July 9, 2026
Last Updated
August 27, 2026

Compare GPT-5.6 Sol

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.