GPT-5.6 Luna vs GPT-5.6 Sol: Specs & Benchmark Comparison
| Characteristic | GPT-5.6 Luna | GPT-5.6 Sol |
|---|---|---|
| Company | OpenAI | OpenAI |
| Release Date | July 9, 2026 | July 9, 2026 |
| Parameters | — | — |
| Multimodal | Yes | Yes |
| Context (input) | 1.1M | 1.1M |
| Context (output) | 128K | 128K |
| Input Price / 1M | $0.20 | $5.00 |
| Output Price / 1M | $1.20 | $30.00 |
| Average Score | 0.5 | 0.6 |
| Benchmarks | ||
| MRCR v2 (8-needle) | 0.4 | 0.9 |
| ExploitBench | 0.3 | 0.7 |
| KernelGen 1P | 0.2 | 0.6 |
| MRCR v2 (8-needle, 512K-1M) | 0.4 | 0.7 |
| Graphwalks BFS 1M | 0.5 | 0.8 |
| FrontierMath Tier 4 (v2) | 0.6 | 0.8 |
| SEC-bench Pro | 0.5 | 0.7 |
| ExploitGym | 0.1 | 0.3 |
| PostTrainBench Lite | 0.3 | 0.5 |
| GeneBench-Pro | 0.1 | 0.3 |
| MedChemBench (Internal) | 0.3 | 0.5 |
| Internal Research Debugging Evaluation | 0.5 | 0.7 |
| Big Finance Bench | 0.4 | 0.5 |
| OSWorld 2.0 | 0.5 | 0.6 |
| RSI Index | 0.4 | 0.6 |
| Capture-the-Flag Challenges (Internal) | 0.9 | 1.0 |
| FrontierMath | 0.8 | 0.9 |
| BenchCAD (with Python tool) | 0.7 | 0.8 |
| Graphwalks BFS >128k | 0.8 | 0.9 |
| LifeSciBench | 0.5 | 0.6 |
| NanoGPT | 0.0 | 0.1 |
| GDP.pdf | 0.2 | 0.3 |
| Artificial Analysis | 0.5 | 0.6 |
| Management Consulting Tasks (Internal) | 0.4 | 0.4 |
| FrontierCode 1.1 | 0.4 | 0.5 |
| ARC-AGI-3 | 0.0 | 0.1 |
| BenchCAD | 0.6 | 0.7 |
| BrowseComp | 0.8 | 0.9 |
| DeepSWE 1.1 | 0.7 | 0.7 |
| DeepSWE | 0.7 | 0.7 |
| MMMU-Pro (with tools) | 0.8 | 0.8 |
| HealthBench Professional | 0.6 | 0.6 |
| MMMU-Pro | 0.8 | 0.8 |
| Toolathlon | 0.5 | 0.6 |
| Terminal-Bench 2.1 | 0.8 | 0.9 |
| AutomationBench | 0.1 | 0.2 |
| Agents' Last Exam | 0.5 | 0.5 |
| GPQA | 0.9 | 0.9 |
| SWE-Bench Pro | 0.6 | 0.6 |
| Search and Function-Calling | 0.9 | 0.9 |
| HealthBench | 0.6 | 0.6 |
| HealthBench Hard | 0.3 | 0.3 |
| HealthBench Consensus | 1.0 | 1.0 |
| Connectors | 1.0 | 1.0 |
Visual Benchmark Comparison
GPT-5.6 Luna
GPT-5.6 Sol
MRCR v2 (8-needle)0.4 vs 0.9
0.4
0.9
ExploitBench0.3 vs 0.7
0.3
0.7
KernelGen 1P0.2 vs 0.6
0.2
0.6
MRCR v2 (8-needle, 512K-1M)0.4 vs 0.7
0.4
0.7
Graphwalks BFS 1M0.5 vs 0.8
0.5
0.8
FrontierMath Tier 4 (v2)0.6 vs 0.8
0.6
0.8
SEC-bench Pro0.5 vs 0.7
0.5
0.7
ExploitGym0.1 vs 0.3
0.1
0.3
PostTrainBench Lite0.3 vs 0.5
0.3
0.5
GeneBench-Pro0.1 vs 0.3
0.1
0.3
MedChemBench (Internal)0.3 vs 0.5
0.3
0.5
Internal Research Debugging Evaluation0.5 vs 0.7
0.5
0.7
Big Finance Bench0.4 vs 0.5
0.4
0.5
OSWorld 2.00.5 vs 0.6
0.5
0.6
RSI Index0.4 vs 0.6
0.4
0.6
Verdict
GPT-5.6 Luna leads in 1 out of 4 comparison categories.
Overall Performance
Both models show comparable average scores: GPT-5.6 Luna — 0.5, GPT-5.6 Sol — 0.6.
API Cost
GPT-5.6 Luna is 25.0x cheaper: input $0.20/1M vs $5.00/1M tokens.
Context Window
Same context size: 1M tokens.
Recency
Both models were released around the same time: 7/9/2026 and 7/9/2026.
More About These Models
Related Comparisons
Frequently Asked Questions
Which is better for coding — GPT-5.6 Luna or GPT-5.6 Sol?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — GPT-5.6 Luna or GPT-5.6 Sol?
GPT-5.6 Luna is cheaper for input: $0.20 per 1M tokens vs $5.00.
Which has a larger context window — GPT-5.6 Luna or GPT-5.6 Sol?
GPT-5.6 Luna supports a larger context: 1,050,000 tokens vs 1,050,000.
The GPT-5.6 Luna and GPT-5.6 Sol comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the GPT-5.6 Luna or GPT-5.6 Sol page. See also the complete list of AI model comparisons.