GPT-5.2
MultimodalGPT-5.2 demonstrates significant improvements in professional tasks, outperforming experts on GDPval with 70.9% wins or ties. Sets new records in coding (SWE-Bench Pro 55.6%), science (GPQA Diamond ~92-93%), math (AIME 2025: 100%), long-context accuracy up to 256k tokens, and reliable tool calling (Tau2 Telecom 98.7%). Available in Instant, Thinking, and Pro variants.
Key Specifications
Parameters
-
Context
400.0K
Release Date
December 10, 2025
Average Score
78.2%
Timeline
Key dates in the model's history
Announcement
December 10, 2025
Today
October 5, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
August 1, 2025
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$1.75
Output (per 1M tokens)
$14.00
Max Input Tokens
400.0K
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
GPT-5.2 Thinking - SWE-bench Verified. • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
GPT-5.2 Thinking - GPQA (without tools) • Self-reported
Other Tests
Specialized benchmarks
AIME 2025
GPT-5.2 Thinking - AIME 2025 (without tools) • Self-reported
TAU2 Telecom
GPT-5.2 Thinking - Tau2-bench • Self-reported
HMMT 2025
GPT-5.2 Thinking - HMMT 2025 (without tools) • Self-reported
ARC-AGI
GPT-5.2 Thinking - ARC-AGI-1 (Verified). • Self-reported
ARC-AGI v2
GPT-5.2 Thinking - ARC-AGI-2 (Verified). • Self-reported
BrowseComp
GPT-5.2 Thinking - BrowseComp. • Self-reported
BrowseComp Long Context 128k
GPT-5.2 Thinking - BrowseComp Long Context 128k. • Self-reported
BrowseComp Long Context 256k
GPT-5.2 Thinking - BrowseComp Long Context 256k. • Self-reported
CharXiv-R
GPT-5.2 Thinking - CharXiv reasoning (no tools). • Self-reported
FrontierMath
GPT-5.2 Thinking - FrontierMath Tier 1-3 (with Python). • Self-reported
Graphwalks BFS <128k
GPT-5.2 Thinking - GraphWalks BFS <128k. • Self-reported
Graphwalks parents <128k
GPT-5.2 Thinking - GraphWalks parents <128k. • Self-reported
Humanity's Last Exam
GPT-5.2 Thinking - Humanity's Last Exam (no tools). • Self-reported
LiveBench
2026-01-08, High • Verified
MCP Atlas
GPT-5.2 Thinking - Scale MCP-Atlas. • Self-reported
MMMLU
GPT-5.2 Thinking - Multilingual MMLU. • Self-reported
MMMU-Pro
GPT-5.2 Thinking - MMMU Pro (no tools). • Self-reported
ScreenSpot Pro
GPT-5.2 Thinking - ScreenSpot Pro (with Python). • Self-reported
SWE-Lancer (IC-Diamond subset)
GPT-5.2 Thinking - SWE-Lancer IC Diamond subset. • Self-reported
Tau2 Retail
GPT-5.2 Thinking - Tau2-bench Retail. • Self-reported
Toolathlon
GPT-5.2 Thinking - Tool Decathlon. • Self-reported
VideoMMMU
GPT-5.2 Thinking - Video MMMU (no tools). • Self-reported
License & Metadata
License
proprietary
Announcement Date
December 10, 2025
Last Updated
December 10, 2025
Similar Models
All ModelsGPT-5.2 Pro
OpenAI
MM
Best score:0.9 (GPQA)
Released:Dec 2025
Price:$21.00/1M tokens
GPT-6 Astra
OpenAI
MM
Best score:1.0 (GPQA)
Released:Sep 2026
Price:$10.00/1M tokens
GPT-5.4
OpenAI
MM
Best score:1.0 (TAU)
Released:Mar 2026
Price:$2.50/1M tokens
GPT-5.5
OpenAI
MM
Best score:1.0 (TAU)
Released:Apr 2026
Price:$5.00/1M tokens
GPT-5.6 Sol
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$5.00/1M tokens
GPT-5.4 mini
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.75/1M tokens
GPT-5.6 Terra
OpenAI
MM
Best score:0.9 (GPQA)
Released:Jul 2026
Price:$2.00/1M tokens
GPT-5.4 nano
OpenAI
MM
Best score:0.9 (TAU)
Released:Mar 2026
Price:$0.20/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.