Key Specifications
Parameters
-
Context
1.1M
Release Date
July 9, 2026
Average Score
52.3%
Timeline
Key dates in the model's history
Announcement
July 9, 2026
Today
August 27, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
February 16, 2026
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.20
Output (per 1M tokens)
$1.20
Max Input Tokens
1.1M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Reasoning
Logical reasoning and analysis
GPQA
GPQA Diamond. Max reasoning effort. • Self-reported
Other Tests
Specialized benchmarks
Agents' Last Exam
Agents' Last Exam. Max reasoning effort. • Self-reported
ARC-AGI-3
ARC-AGI-3. Max reasoning effort. • Self-reported
Artificial Analysis
Artificial Analysis Intelligence Index v4.1 (max). Score 51. • Self-reported
AutomationBench
AutomationBench. Max reasoning effort. • Self-reported
BenchCAD
BenchCAD (no tools). Max reasoning effort. • Self-reported
BenchCAD (with Python tool)
BenchCAD with Python tool. Max reasoning effort. • Self-reported
Big Finance Bench
Big Finance Bench. Max reasoning effort. • Self-reported
BrowseComp
BrowseComp. Max reasoning effort. • Self-reported
Capture-the-Flag Challenges (Internal)
Capture-the-Flag Challenges (Internal). Max reasoning effort. • Self-reported
Connectors
Connectors production benchmark pass rate. • Self-reported
DeepSWE
DeepSWE v1.1. Max reasoning effort. • Self-reported
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 67% ± 4%. • Self-reported
ExploitBench
ExploitBench. Max reasoning effort. • Self-reported
ExploitGym
ExploitGym, six-hour cap. Max reasoning effort. • Self-reported
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at max effort: 39.8%. • Self-reported
FrontierMath
FrontierMath Tier 1-3 (v2). Max reasoning effort. • Self-reported
FrontierMath Tier 4 (v2)
FrontierMath Tier 4 (v2). Max reasoning effort. • Self-reported
GDP.pdf
gdp.pdf. Max reasoning effort. • Self-reported
GeneBench-Pro
GeneBench Pro. Max reasoning effort. • Self-reported
Graphwalks BFS >128k
GraphWalks BFS 256k f1. Max reasoning effort. • Self-reported
Graphwalks BFS 1M
GraphWalks BFS 1M f1. Max reasoning effort. • Self-reported
HealthBench
HealthBench, length-adjusted score (unadjusted 55.4). • Self-reported
HealthBench Consensus
HealthBench Consensus, length-adjusted score (unadjusted 95.1). • Self-reported
HealthBench Hard
HealthBench Hard, length-adjusted score (unadjusted 31.4). • Self-reported
HealthBench Professional
HealthBench Professional, length-adjusted score (unadjusted 59.8). • Self-reported
Internal Research Debugging Evaluation
Internal Research Debugging Evaluation. Max reasoning effort. • Self-reported
KernelGen 1P
KernelGen 1P. Max reasoning effort. • Self-reported
LifeSciBench
LifeSciBench. Max reasoning effort. • Self-reported
Management Consulting Tasks (Internal)
Management Consulting Tasks (Internal). Max reasoning effort. • Self-reported
MedChemBench (Internal)
MedChemBench (Internal). Max reasoning effort. • Self-reported
MMMU-Pro
MMMU Pro (no tools). Max reasoning effort. • Self-reported
MMMU-Pro (with tools)
MMMU Pro (with tools). Max reasoning effort. • Self-reported
MRCR v2 (8-needle)
OpenAI MRCR v2 8-needle, 256K-512K. Max reasoning effort. • Self-reported
MRCR v2 (8-needle, 512K-1M)
OpenAI MRCR v2 8-needle, 512K-1M. Max reasoning effort. • Self-reported
NanoGPT
NanoGPT. Max reasoning effort. • Self-reported
OSWorld 2.0
OSWorld 2.0 binary completion. Max reasoning effort. • Self-reported
PostTrainBench Lite
PostTrainBench Lite mean reward. Max reasoning effort. • Self-reported
RSI Index
RSI Index (aggregate self-improvement). Max reasoning effort. • Self-reported
Search and Function-Calling
Search and Function-Calling production benchmark pass rate. • Self-reported
SEC-bench Pro
SEC-Bench Pro. Max reasoning effort. • Self-reported
SWE-Bench Pro
SWE-Bench Pro. Max reasoning effort. • Self-reported
Terminal-Bench 2.1
Terminal-Bench 2.1. Max reasoning effort. • Self-reported
Toolathlon
Toolathlon. Max reasoning effort. • Self-reported
License & Metadata
License
proprietary
Announcement Date
July 9, 2026
Last Updated
August 27, 2026
Compare GPT-5.6 Luna
All comparisonsSimilar Models
All ModelsGPT-5
OpenAI
MM
Best score:0.9 (HumanEval)
Released:Aug 2025
Price:$1.25/1M tokens
GPT-5.5
OpenAI
MM
Best score:1.0 (TAU)
Released:Apr 2026
Price:$5.00/1M tokens
GPT-5.6 Terra
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$2.00/1M tokens
GPT-5.1 Instant
OpenAI
MM
Best score:1.0 (TAU)
Released:Nov 2025
Price:$0.30/1M tokens
GPT-5.4 mini
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.42/1M tokens
GPT-5.4 nano
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.12/1M tokens
GPT-5.5 Instant
OpenAI
MM
Best score:0.9 (GPQA)
Released:May 2026
Price:$5.00/1M tokens
GPT-5.6 Sol
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$5.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.