GPT OSS 120B vs GPT OSS 20B: Specs & Benchmark Comparison

GPT OSS 120B and GPT OSS 20B both come from OpenAI, and both were released in August 2025. GPT OSS 120B is the larger model at roughly 120 billion parameters, against 20 billion for GPT OSS 20B.

The two models share 9 published benchmarks. GPT OSS 120B leads on 8 of them, GPT OSS 20B on 1. The widest gaps are on Codeforces Competition code, where GPT OSS 20B scores 74.3% against 26.2%; HealthBench Hard - Challenging health conversations, where GPT OSS 120B scores 30.0% against 10.8%. Averaged across everything we track, GPT OSS 120B sits at 52.0% and GPT OSS 20B at 43.6%.

GPT OSS 20B is the cheaper API at $0.10 per million input tokens and $0.50 per million output tokens, about 50% below GPT OSS 120B at $0.15 and $0.60. Both accept a context window of about 131K tokens.

CharacteristicGPT OSS 120BGPT OSS 20B
CompanyOpenAIOpenAI
Release DateAugust 5, 2025August 5, 2025
Parameters120B20B
MultimodalYesYes
Context (input)131K131K
Context (output)30K30K
Input Price / 1M$0.15$0.10
Output Price / 1M$0.60$0.50
Average Score52.0%43.6%
Benchmarks
Codeforces Competition code26.2%74.3%
HealthBench Hard - Challenging health conversations30.0%10.8%
HealthBench - Realistic health conversations57.6%42.5%
TAU-bench Retail benchmark67.8%54.8%
GPQA80.1%71.5%
Humanity's Last Exam19.0%10.9%
Codeforces Competition code82.1%74.3%
MMLU benchmark90.0%85.3%
Humanity's Last Exam14.9%10.9%

Visual Benchmark Comparison

GPT OSS 120B
GPT OSS 20B
Codeforces Competition code0.3 vs 0.7
0.3
0.7
HealthBench Hard - Challenging health conversations0.3 vs 0.1
0.3
0.1
HealthBench - Realistic health conversations0.6 vs 0.4
0.6
0.4
TAU-bench Retail benchmark0.7 vs 0.5
0.7
0.5
GPQA0.8 vs 0.7
0.8
0.7
Humanity's Last Exam0.2 vs 0.1
0.2
0.1
Codeforces Competition code0.8 vs 0.7
0.8
0.7
MMLU benchmark0.9 vs 0.9
0.9
0.9
Humanity's Last Exam0.1 vs 0.1
0.1
0.1

Verdict

GPT OSS 20B leads in 1 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: GPT OSS 120B — 0.5, GPT OSS 20B — 0.4.

API Cost

GPT OSS 20B is 1.3x cheaper: input $0.10/1M vs $0.15/1M tokens.

Context Window

Same context size: 131K tokens.

Recency

Both models were released around the same time: 8/5/2025 and 8/5/2025.

More About These Models

Frequently Asked Questions

Which is better for coding — GPT OSS 120B or GPT OSS 20B?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — GPT OSS 120B or GPT OSS 20B?
GPT OSS 20B is cheaper for input: $0.10 per 1M tokens vs $0.15.
Which has a larger context window — GPT OSS 120B or GPT OSS 20B?
GPT OSS 120B supports a larger context: 131,000 tokens vs 131,000.

The GPT OSS 120B and GPT OSS 20B comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the GPT OSS 120B or GPT OSS 20B page. See also the complete list of AI model comparisons.