DeepSeek-V4-Flash-0731 vs GPT-5.6 Luna: Specs & Benchmark Comparison

DeepSeek-V4-Flash-0731 is developed by DeepSeek, while GPT-5.6 Luna comes from OpenAI. Both were released in July 2026. DeepSeek-V4-Flash-0731 has a published size of about 304 billion parameters; OpenAI has not disclosed the parameter count of GPT-5.6 Luna.

The two models share 5 published benchmarks. GPT-5.6 Luna leads on 3 of them, DeepSeek-V4-Flash-0731 on 2. The widest gaps are on Agents' Last Exam, where GPT-5.6 Luna scores 50.3% against 25.2%; Toolathlon, where DeepSeek-V4-Flash-0731 scores 70.0% against 53.4%. Averaged across everything we track, DeepSeek-V4-Flash-0731 sits at 57.5% and GPT-5.6 Luna at 51.5%.

DeepSeek-V4-Flash-0731 is the cheaper API at $0.14 per million input tokens and $0.28 per million output tokens, about 43% below GPT-5.6 Luna at $0.20 and $1.20. Both accept a context window of about 1M tokens. DeepSeek-V4-Flash-0731 accepts text as input, while GPT-5.6 Luna accepts text and images.

CharacteristicDeepSeek-V4-Flash-0731GPT-5.6 Luna
CompanyDeepSeekOpenAI
Release DateJuly 31, 2026July 9, 2026
Parameters304B
MultimodalNoYes
Context (input)1.0M1.1M
Context (output)393K128K
Input Price / 1M$0.14$0.20
Output Price / 1M$0.28$1.20
Average Score57.5%51.5%
Benchmarks
Agents' Last Exam25.2%50.3%
Toolathlon70.0%53.4%
DeepSWE54.4%67.2%
AutomationBench25.1%14.9%
Terminal-Bench 2.183.0%84.7%

Visual Benchmark Comparison

DeepSeek-V4-Flash-0731
GPT-5.6 Luna
Agents' Last Exam0.3 vs 0.5
0.3
0.5
Toolathlon0.7 vs 0.5
0.7
0.5
DeepSWE0.5 vs 0.7
0.5
0.7
AutomationBench0.3 vs 0.1
0.3
0.1
Terminal-Bench 2.10.8 vs 0.8
0.8
0.8

Verdict

DeepSeek-V4-Flash-0731 leads in 2 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: DeepSeek-V4-Flash-0731 — 0.6, GPT-5.6 Luna — 0.5.

API Cost

DeepSeek-V4-Flash-0731 is 3.3x cheaper: input $0.14/1M vs $0.20/1M tokens.

Context Window

GPT-5.6 Luna supports a larger context: 1M vs 1M tokens.

Recency

DeepSeek-V4-Flash-0731 is newer: released 7/31/2026 vs 7/9/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — DeepSeek-V4-Flash-0731 or GPT-5.6 Luna?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — DeepSeek-V4-Flash-0731 or GPT-5.6 Luna?
DeepSeek-V4-Flash-0731 is cheaper for input: $0.14 per 1M tokens vs $0.20.
Which has a larger context window — DeepSeek-V4-Flash-0731 or GPT-5.6 Luna?
GPT-5.6 Luna supports a larger context: 1,050,000 tokens vs 1,000,000.

The DeepSeek-V4-Flash-0731 and GPT-5.6 Luna comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the DeepSeek-V4-Flash-0731 or GPT-5.6 Luna page. See also the complete list of AI model comparisons.