Key Specifications
Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
34.2%
Timeline
Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
September 10, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
May 31, 2024
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.10
Output (per 1M tokens)
$0.40
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
Standard benchmark
AI: Alright, I'll solve this step-by-step. • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Diamond • Self-reported
Multimodal
Working with images and visual data
MathVista
Standard benchmark • Self-reported
MMMU
Standard benchmark • Self-reported
Other Tests
Specialized benchmarks
AIME 2024
Standard benchmark AI: I on question, using its capabilities, and I reasoning in answer. Human: [] • Self-reported
IFEval
Standard benchmark
AI: Translate following text • Self-reported
Aider-Polyglot
Standard benchmark
AI: I'm sorry, but your request is unclear. Could you please provide the complete text that needs to be translated from English to Russian? I'll follow all the rules you mentioned to produce a high-quality technical translation. • Self-reported
MultiChallenge
Standard benchmark (GPT-4o grader) • Self-reported
Aider-Polyglot Edit
Standard benchmark
Standard benchmark
AI: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, ... • Self-reported
MMMLU
Standard benchmark • Self-reported
Multi-IF
Standard benchmark • Self-reported
TAU-bench Retail
Average value by 5 without use tools/prompts ([4], model GPT-4o) • Self-reported
TAU-bench Airline
Average from 5 without tools/prompts ([4]) • Self-reported
CharXiv-R
Standard benchmark • Self-reported
Internal API instruction following (hard)
Internal benchmark • Self-reported
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini grader, [3]) • Self-reported
COLLIE
Standard benchmark AI: Translate this text fully, full translation • Self-reported
OpenAI-MRCR: 2 needle 128k
Internal benchmark AI: • Self-reported
OpenAI-MRCR: 2 needle 1M
Internal benchmark
AI:
Internal benchmark • Self-reported
Graphwalks BFS <128k
Standard benchmark AI: Using knowledge and abilities, I, you solve following tasks: evaluation: For each question full solution. intermediate steps, course reasoning and final answer. Task: [] • Self-reported
Graphwalks BFS >128k
Internal benchmark • Self-reported
Graphwalks parents <128k
Internal benchmark • Self-reported
Graphwalks parents >128k
Internal benchmark AI: to query, translation • Self-reported
CharXiv-D
Standard benchmark
Standard benchmark
AI: 1 • Self-reported
ComplexFuncBench
Standard benchmark • Self-reported
License & Metadata
License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025
Similar Models
All ModelsGPT-5.1 Codex Mini
OpenAI
MM
Released:Nov 2025
Price:$0.25/1M tokens
GPT-5.3 Chat
OpenAI
MM
Released:Mar 2026
Price:$1.75/1M tokens
GPT-5.4 Pro
OpenAI
MM
Released:Mar 2026
Price:$15.00/1M tokens
o3-pro
OpenAI
MM
Released:Jun 2025
Price:$20.00/1M tokens
GPT-5.1 Medium
OpenAI
MM
Released:Nov 2025
Price:$1.25/1M tokens
GPT-5.1 Codex High
OpenAI
MM
Released:Nov 2025
Price:$1.25/1M tokens
GPT-5.3 Codex
OpenAI
MM
Released:Feb 2026
Price:$1.75/1M tokens
GPT-5.2 Codex
OpenAI
MM
Released:Jan 2026
Price:$1.75/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.