OpenAI logo

GPT-5.1 Thinking

Multimodal
OpenAI

GPT-5.1 Thinking is the next iteration of GPT-5 with improved adaptive reasoning capabilities. Unlike GPT-5.1 Instant, GPT-5.1 Thinking more precisely adapts thinking time to each question, providing deeper reasoning for complex queries. The model demonstrates significant improvements in coding and complex reasoning scenarios.

Key Specifications

Parameters
-
Context
400.0K
Release Date
November 11, 2025
Average Score
77.9%

Timeline

Key dates in the model's history
Announcement
November 11, 2025
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.25
Output (per 1M tokens)
$10.00
Max Input Tokens
400.0K
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
GPT-5.1 Thinking with adaptive reasoning - SWE-bench Verified (all 500 problems). • Self-reported
76.3%

Reasoning

Logical reasoning and analysis
GPQA
GPT-5.1 Thinking - GPQA Diamond (without tools) • Self-reported
88.0%

Multimodal

Working with images and visual data
MMMU
GPT-5.1 Thinking with tasks level with • Self-reported
85.0%

Other Tests

Specialized benchmarks
TAU2 Telecom
GPT-5.1 Thinking with Benchmark functions () • Self-reported
96.0%
AIME 2025
GPT-5.1 Thinking with (without tools) - AIME 2025 mathematics • Self-reported
94.0%
BrowseComp Long Context 128k
GPT-5.1 Thinking with adaptive reasoning - BrowseComp long-context 128k variant. • Self-reported
90.0%
FrontierMath
GPT-5.1 Thinking with adaptive reasoning (with Python tool) - FrontierMath Tier 1-3 expert-level mathematics. • Self-reported
26.7%
Tau2 Airline
GPT-5.1 Thinking with adaptive reasoning - Function calling benchmark (airline domain). • Self-reported
67.0%
Tau2 Retail
GPT-5.1 Thinking with adaptive reasoning - Function calling benchmark (retail domain). • Self-reported
77.9%

License & Metadata

License
proprietary
Announcement Date
November 11, 2025
Last Updated
November 11, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.