GPT-5.6 Luna vs Qwen3.8 Max: Specs & Benchmark Comparison

GPT-5.6 Luna is developed by OpenAI, while Qwen3.8 Max comes from Alibaba. GPT-5.6 Luna was released in July 2026, and Qwen3.8 Max followed a month later in August 2026. Qwen3.8 Max has a published size of about 2.4 trillion parameters; OpenAI has not disclosed the parameter count of GPT-5.6 Luna.

The two models share 10 published benchmarks. Qwen3.8 Max leads on 9 of them, GPT-5.6 Luna on 1. The widest gaps are on MRCR v2 (8-needle), where Qwen3.8 Max scores 93.0% against 41.3%; Toolathlon, where Qwen3.8 Max scores 72.5% against 53.4%. Averaged across everything we track, GPT-5.6 Luna sits at 51.5% and Qwen3.8 Max at 71.0%.

GPT-5.6 Luna is the cheaper API at $0.20 per million input tokens and $1.20 per million output tokens, roughly 13 times cheaper than Qwen3.8 Max at $2.50 and $6.25. Both accept a context window of about 1M tokens. Both accept text and images as input.

CharacteristicGPT-5.6 LunaQwen3.8 Max
CompanyOpenAIAlibaba
Release DateJuly 9, 2026August 2, 2026
Parameters2.4T
MultimodalYesYes
Context (input)1.1M1.0M
Context (output)128K131K
Input Price / 1M$0.20$2.50
Output Price / 1M$1.20$6.25
Average Score51.5%71.0%
Benchmarks
MRCR v2 (8-needle)41.3%93.0%
Toolathlon53.4%72.5%
AutomationBench14.9%27.3%
DeepSWE 1.167.0%56.6%
SWE-Bench Pro62.7%67.7%
HealthBench55.8%60.2%
MMMU-Pro78.4%82.3%
Agents' Last Exam50.3%52.4%
Terminal-Bench 2.184.7%86.6%
GPQA92.3%92.6%

Visual Benchmark Comparison

GPT-5.6 Luna
Qwen3.8 Max
MRCR v2 (8-needle)0.4 vs 0.9
0.4
0.9
Toolathlon0.5 vs 0.7
0.5
0.7
AutomationBench0.1 vs 0.3
0.1
0.3
DeepSWE 1.10.7 vs 0.6
0.7
0.6
SWE-Bench Pro0.6 vs 0.7
0.6
0.7
HealthBench0.6 vs 0.6
0.6
0.6
MMMU-Pro0.8 vs 0.8
0.8
0.8
Agents' Last Exam0.5 vs 0.5
0.5
0.5
Terminal-Bench 2.10.8 vs 0.9
0.8
0.9
GPQA0.9 vs 0.9
0.9
0.9

Verdict

GPT-5.6 Luna leads in 2 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: GPT-5.6 Luna — 0.5, Qwen3.8 Max — 0.7.

API Cost

GPT-5.6 Luna is 6.3x cheaper: input $0.20/1M vs $2.50/1M tokens.

Context Window

GPT-5.6 Luna supports a larger context: 1M vs 1M tokens.

Recency

Qwen3.8 Max is newer: released 8/2/2026 vs 7/9/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — GPT-5.6 Luna or Qwen3.8 Max?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — GPT-5.6 Luna or Qwen3.8 Max?
GPT-5.6 Luna is cheaper for input: $0.20 per 1M tokens vs $2.50.
Which has a larger context window — GPT-5.6 Luna or Qwen3.8 Max?
GPT-5.6 Luna supports a larger context: 1,050,000 tokens vs 1,000,000.

The GPT-5.6 Luna and Qwen3.8 Max comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the GPT-5.6 Luna or Qwen3.8 Max page. See also the complete list of AI model comparisons.