Key Specifications
Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
49.6%
Timeline
Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
September 10, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
May 31, 2024
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.40
Output (per 1M tokens)
$1.60
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
Standard benchmark AI: whether more complex this tasks? • Self-reported
Programming
Programming skills tests
SWE-Bench Verified
methodology, on [2] • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Diamond • Self-reported
Multimodal
Working with images and visual data
MathVista
Standard benchmark
AI: Human_Evaluation • Self-reported
MMMU
Standard benchmark • Self-reported
Other Tests
Specialized benchmarks
AIME 2024
Standard benchmark
AI: I'll solve this step-by-step. • Self-reported
IFEval
Standard benchmark • Self-reported
Aider-Polyglot
Standard benchmark • Self-reported
MultiChallenge
Standard benchmark (GPT-4o grader) • Self-reported
Aider-Polyglot Edit
Standard benchmark
AI: (1/8) Standard benchmark • Self-reported
MMMLU
Standard benchmark
AI: (7 of 25 marks) • Self-reported
Multi-IF
Standard benchmark
AI: I'm ChatGPT, a large language model trained by OpenAI, based on the GPT-4 architecture. • Self-reported
TAU-bench Retail
Average value by 5 without tools/prompts (note [4], model GPT-4o) • Self-reported
TAU-bench Airline
Average by 5 without use special tools/prompts ([4]) • Self-reported
CharXiv-R
Standard benchmark • Self-reported
Internal API instruction following (hard)
Internal benchmark AI: *internal thoughts* This too text for translation. I its exactly Internal benchmark • Self-reported
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini, [3]) • Self-reported
COLLIE
Standard benchmark AI: (thinking) Standard benchmark = standard benchmark. "benchmark" without translation, so how this in AI • Self-reported
OpenAI-MRCR: 2 needle 128k
Internal benchmark • Self-reported
OpenAI-MRCR: 2 needle 1M
Internal benchmark • Self-reported
Graphwalks BFS <128k
Standard benchmark
Standard benchmark
AI • Self-reported
Graphwalks BFS >128k
Internal benchmark • Self-reported
Graphwalks parents <128k
Internal benchmark • Self-reported
Graphwalks parents >128k
Internal benchmark • Self-reported
CharXiv-D
Standard benchmark AI: on this task. First let's solve her/its method. Task: $x$ such, that $2^x = 32$. In order to find $x$, I I can use $2^x = 32$ $2^x = 2^5$ (so how $32 = 2^5$) should scores, therefore $x = 5$. Answer: $x = 5$ • Self-reported
ComplexFuncBench
Standard benchmark • Self-reported
AIME 2025
GPT-4.1 mini without tools - mathematics (AIME 2025) • Self-reported
Humanity's Last Exam
GPT-4.1 mini without tools - Questions expert level by various subjects. • Self-reported
HMMT 2025
GPT-4.1 mini without tools - Harvard-MIT Mathematics Tournament. • Self-reported
License & Metadata
License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025
Similar Models
All Modelso4-mini
OpenAI
MM
Best score:0.8 (GPQA)
Released:Apr 2025
Price:$1.10/1M tokens
GPT-4o
OpenAI
MM
Best score:0.9 (MMLU)
Released:Aug 2024
Price:$2.50/1M tokens
GPT-5 nano
OpenAI
MM
Best score:0.7 (GPQA)
Released:Aug 2025
Price:$0.05/1M tokens
GPT-4
OpenAI
MM
Best score:1.0 (ARC)
Released:Jun 2023
Price:$30.00/1M tokens
GPT-4o mini
OpenAI
MM
Best score:0.9 (HumanEval)
Released:Jul 2024
Price:$0.15/1M tokens
GPT-5.5 Pro
OpenAI
MM
Released:Apr 2026
Price:$30.00/1M tokens
o3-pro
OpenAI
MM
Released:Jun 2025
Price:$20.00/1M tokens
GPT-5.1 Medium
OpenAI
MM
Released:Nov 2025
Price:$1.25/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.