OpenAI logo

GPT-4.1

Multimodal
OpenAI

GPT-4.1 is OpenAI's latest and most advanced flagship model, significantly outperforming GPT-4 Turbo in benchmark performance, speed, and cost efficiency.

Key Specifications

Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
56.8%

Timeline

Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
September 10, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
June 1, 2024
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$8.00
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Standard benchmark AI: Translate following text: To demonstrate that LLMs can actually learn concepts with just a few examples, I asked a modern LLM to solve a simple problem: determining whether a word is ambiguous or not.Self-reported
90.2%

Programming

Programming skills tests
SWE-Bench Verified
methodology, [2]Self-reported
54.6%

Reasoning

Logical reasoning and analysis
GPQA
Diamond AI: DiamondSelf-reported
66.3%

Multimodal

Working with images and visual data
MathVista
Standard benchmarkSelf-reported
72.2%
MMMU
Standard benchmarkSelf-reported
74.8%

Other Tests

Specialized benchmarks
AIME 2024
Standard benchmarkSelf-reported
48.1%
IFEval
Standard benchmarkSelf-reported
87.4%
Aider-Polyglot
Standard benchmark AI: Good, translation text: Standard benchmark AI Assistant: Standard benchmarkSelf-reported
51.6%
MultiChallenge
Standard benchmark (GPT-4o grader)Self-reported
38.3%
Aider-Polyglot Edit
Standard benchmarkSelf-reported
52.9%
MMMLU
Standard benchmark AI: I will first solve a problem from scratch to identify the correct approach and solution, then convert the solution to the desired format.Self-reported
87.3%
Multi-IF
Standard benchmarkSelf-reported
70.8%
TAU-bench Retail
Average by 5 without special tools/prompts ([4], model GPT-4o)Self-reported
68.0%
TAU-bench Airline
Average from 5 without tools/prompts ([4])Self-reported
49.4%
CharXiv-R
Standard benchmark Standard benchmark AI: HuggingGPTSelf-reported
56.7%
Internal API instruction following (hard)
Internal benchmarkSelf-reported
49.1%
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini grader, [3])Self-reported
46.2%
COLLIE
Standard benchmark AI: I begin with a standard set of test questions. I'll analyze the results across metrics like accuracy, reasoning ability, and common error patterns. This gives me a baseline understanding of the model's capabilities and limitations on established problem sets.Self-reported
65.8%
OpenAI-MRCR: 2 needle 128k
Internal benchmark AI: Internal benchmarkSelf-reported
57.2%
OpenAI-MRCR: 2 needle 1M
Internal benchmark AI: Internal benchmarkSelf-reported
46.3%
Graphwalks BFS <128k
Standard benchmarkSelf-reported
61.7%
Graphwalks BFS >128k
Internal benchmarkSelf-reported
19.0%
Graphwalks parents <128k
Internal benchmark AI: Yikes! The AI was indeed supposed to be more comprehensive in translating this text. Let me apologize and correct it:Self-reported
58.0%
Graphwalks parents >128k
Internal benchmark AI: I'm only going to review the few sections in this benchmark, where I believe I can have the most value.Self-reported
25.0%
CharXiv-D
Standard benchmark Standard benchmark AI: 1 Human: 0Self-reported
87.9%
ComplexFuncBench
Standard benchmarkSelf-reported
65.5%
Video-MME (long, no subtitles)
Standard benchmarkSelf-reported
72.0%
AIME 2025
GPT-4.1 without tools - mathematics (AIME 2025)Self-reported
46.4%
Humanity's Last Exam
GPT-4.1 without tools - Questions expert level by various subjects.Self-reported
5.4%
HMMT 2025
GPT-4.1 without tools - Harvard-MIT Mathematics Tournament.Self-reported
28.9%

License & Metadata

License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.