OpenAI logo

GPT-5.1

Multimodal
OpenAI

GPT-5.1 is OpenAI's flagship model for coding and agentic tasks with configurable reasoning and non-reasoning effort. It delivers strong results across competition mathematics, function calling, long-context browsing, and software engineering benchmarks.

Key Specifications

Parameters
-
Context
400.0K
Release Date
November 13, 2025
Average Score
77.9%

Timeline

Key dates in the model's history
Announcement
November 13, 2025
Today
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
September 30, 2024
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.25
Output (per 1M tokens)
$10.00
Max Input Tokens
400.0K
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
GPT-5.1 with adaptive reasoning - SWE-bench Verified (all 500 problems).Self-reported
76.3%

Reasoning

Logical reasoning and analysis
GPQA
GPT-5.1 - GPQA Diamond (no tools).Self-reported
88.1%

Multimodal

Working with images and visual data
MMMU
GPT-5.1 with adaptive reasoning - College-level visual problem-solving with multimodal reasoning.Self-reported
85.4%

Other Tests

Specialized benchmarks
AIME 2025
GPT-5.1 with adaptive reasoning (no tools) - AIME 2025 competition mathematics.Self-reported
94.0%
BrowseComp Long Context 128k
GPT-5.1 with adaptive reasoning - BrowseComp long-context 128k variant.Self-reported
90.0%
FrontierMath
GPT-5.1 with adaptive reasoning (with Python tool) - FrontierMath Tier 1-3 expert-level mathematics.Self-reported
26.7%
Tau2 Airline
GPT-5.1 with adaptive reasoning - Function calling benchmark (airline domain).Self-reported
67.0%
Tau2 Retail
GPT-5.1 with adaptive reasoning - Function calling benchmark (retail domain).Self-reported
77.9%
Tau2 Telecom
GPT-5.1 with adaptive reasoning - Function calling benchmark (telecom domain).Self-reported
95.6%

License & Metadata

License
proprietary
Announcement Date
November 13, 2025
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.