Key Specifications
Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
56.8%
Timeline
Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
September 10, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
June 1, 2024
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$8.00
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
Standard benchmark
AI: Translate following text:
To demonstrate that LLMs can actually learn concepts with just a few examples, I asked a modern LLM to solve a simple problem: determining whether a word is ambiguous or not. • Self-reported
Programming
Programming skills tests
SWE-Bench Verified
methodology, [2] • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Diamond
AI: Diamond • Self-reported
Multimodal
Working with images and visual data
MathVista
Standard benchmark • Self-reported
MMMU
Standard benchmark • Self-reported
Other Tests
Specialized benchmarks
AIME 2024
Standard benchmark • Self-reported
IFEval
Standard benchmark • Self-reported
Aider-Polyglot
Standard benchmark AI: Good, translation text: Standard benchmark AI Assistant: Standard benchmark • Self-reported
MultiChallenge
Standard benchmark (GPT-4o grader) • Self-reported
Aider-Polyglot Edit
Standard benchmark • Self-reported
MMMLU
Standard benchmark
AI: I will first solve a problem from scratch to identify the correct approach and solution, then convert the solution to the desired format. • Self-reported
Multi-IF
Standard benchmark • Self-reported
TAU-bench Retail
Average by 5 without special tools/prompts ([4], model GPT-4o) • Self-reported
TAU-bench Airline
Average from 5 without tools/prompts ([4]) • Self-reported
CharXiv-R
Standard benchmark
Standard benchmark
AI: HuggingGPT • Self-reported
Internal API instruction following (hard)
Internal benchmark • Self-reported
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini grader, [3]) • Self-reported
COLLIE
Standard benchmark
AI: I begin with a standard set of test questions.
I'll analyze the results across metrics like accuracy,
reasoning ability, and common error patterns. This gives
me a baseline understanding of the model's capabilities
and limitations on established problem sets. • Self-reported
OpenAI-MRCR: 2 needle 128k
Internal benchmark
AI: Internal benchmark • Self-reported
OpenAI-MRCR: 2 needle 1M
Internal benchmark
AI: Internal benchmark • Self-reported
Graphwalks BFS <128k
Standard benchmark • Self-reported
Graphwalks BFS >128k
Internal benchmark • Self-reported
Graphwalks parents <128k
Internal benchmark
AI: Yikes! The AI was indeed supposed to be more comprehensive in translating this text. Let me apologize and correct it: • Self-reported
Graphwalks parents >128k
Internal benchmark
AI: I'm only going to review the few sections in this benchmark, where I believe I can have the most value. • Self-reported
CharXiv-D
Standard benchmark
Standard benchmark
AI: 1 Human: 0 • Self-reported
ComplexFuncBench
Standard benchmark • Self-reported
Video-MME (long, no subtitles)
Standard benchmark • Self-reported
AIME 2025
GPT-4.1 without tools - mathematics (AIME 2025) • Self-reported
Humanity's Last Exam
GPT-4.1 without tools - Questions expert level by various subjects. • Self-reported
HMMT 2025
GPT-4.1 without tools - Harvard-MIT Mathematics Tournament. • Self-reported
License & Metadata
License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025
Similar Models
All Modelso4-mini
OpenAI
MM
Best score:0.8 (GPQA)
Released:Apr 2025
Price:$1.10/1M tokens
o3
OpenAI
MM
Best score:0.8 (GPQA)
Released:Apr 2025
Price:$2.00/1M tokens
GPT-4o mini
OpenAI
MM
Best score:0.9 (HumanEval)
Released:Jul 2024
Price:$0.15/1M tokens
GPT-4.5
OpenAI
MM
Best score:0.9 (MMLU)
Released:Feb 2025
Price:$75.00/1M tokens
GPT-4o
OpenAI
MM
Best score:0.9 (MMLU)
Released:Aug 2024
Price:$2.50/1M tokens
GPT-5 nano
OpenAI
MM
Best score:0.7 (GPQA)
Released:Aug 2025
Price:$0.05/1M tokens
GPT-5.1
OpenAI
MM
Best score:0.9 (GPQA)
Released:Nov 2025
Price:$1.25/1M tokens
GPT-4
OpenAI
MM
Best score:1.0 (ARC)
Released:Jun 2023
Price:$30.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.