OpenAI logo

GPT-5.2

Multimodal
OpenAI

GPT-5.2 demonstrates significant improvements in professional tasks, outperforming experts on GDPval with 70.9% wins or ties. Sets new records in coding (SWE-Bench Pro 55.6%), science (GPQA Diamond ~92-93%), math (AIME 2025: 100%), long-context accuracy up to 256k tokens, and reliable tool calling (Tau2 Telecom 98.7%). Available in Instant, Thinking, and Pro variants.

Key Specifications

Parameters
-
Context
400.0K
Release Date
December 10, 2025
Average Score
78.2%

Timeline

Key dates in the model's history
Announcement
December 10, 2025
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
August 1, 2025
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.75
Output (per 1M tokens)
$14.00
Max Input Tokens
400.0K
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
GPT-5.2 Thinking - SWE-bench Verified. • Self-reported
80.0%

Reasoning

Logical reasoning and analysis
GPQA
GPT-5.2 Thinking - GPQA (without tools) • Self-reported
92.0%

Other Tests

Specialized benchmarks
AIME 2025
GPT-5.2 Thinking - AIME 2025 (without tools) • Self-reported
100.0%
TAU2 Telecom
GPT-5.2 Thinking - Tau2-bench • Self-reported
99.0%
HMMT 2025
GPT-5.2 Thinking - HMMT 2025 (without tools) • Self-reported
99.0%
ARC-AGI
GPT-5.2 Thinking - ARC-AGI-1 (Verified). • Self-reported
86.2%
ARC-AGI v2
GPT-5.2 Thinking - ARC-AGI-2 (Verified). • Self-reported
52.9%
BrowseComp
GPT-5.2 Thinking - BrowseComp. • Self-reported
65.8%
BrowseComp Long Context 128k
GPT-5.2 Thinking - BrowseComp Long Context 128k. • Self-reported
92.0%
BrowseComp Long Context 256k
GPT-5.2 Thinking - BrowseComp Long Context 256k. • Self-reported
89.8%
CharXiv-R
GPT-5.2 Thinking - CharXiv reasoning (no tools). • Self-reported
82.1%
FrontierMath
GPT-5.2 Thinking - FrontierMath Tier 1-3 (with Python). • Self-reported
40.3%
Graphwalks BFS <128k
GPT-5.2 Thinking - GraphWalks BFS <128k. • Self-reported
94.0%
Graphwalks parents <128k
GPT-5.2 Thinking - GraphWalks parents <128k. • Self-reported
89.0%
Humanity's Last Exam
GPT-5.2 Thinking - Humanity's Last Exam (no tools). • Self-reported
34.5%
LiveBench
2026-01-08, High • Verified
74.8%
MCP Atlas
GPT-5.2 Thinking - Scale MCP-Atlas. • Self-reported
60.6%
MMMLU
GPT-5.2 Thinking - Multilingual MMLU. • Self-reported
89.6%
MMMU-Pro
GPT-5.2 Thinking - MMMU Pro (no tools). • Self-reported
79.5%
ScreenSpot Pro
GPT-5.2 Thinking - ScreenSpot Pro (with Python). • Self-reported
86.3%
SWE-Lancer (IC-Diamond subset)
GPT-5.2 Thinking - SWE-Lancer IC Diamond subset. • Self-reported
74.6%
Tau2 Retail
GPT-5.2 Thinking - Tau2-bench Retail. • Self-reported
82.0%
Toolathlon
GPT-5.2 Thinking - Tool Decathlon. • Self-reported
46.3%
VideoMMMU
GPT-5.2 Thinking - Video MMMU (no tools). • Self-reported
85.9%

License & Metadata

License
proprietary
Announcement Date
December 10, 2025
Last Updated
December 10, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.