OpenAI logo

GPT-4o

Multimodal
OpenAI

GPT-4o ('o' stands for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in understanding non-English languages, images, and audio.

Key Specifications

Parameters
-
Context
128.0K
Release Date
August 6, 2024
Average Score
52.5%

Timeline

Key dates in the model's history
Announcement
August 6, 2024
Last Update
July 19, 2025
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$10.00
Max Input Tokens
128.0K
Max Output Tokens
16.4K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Accuracy • Self-reported
85.7%

Programming

Programming skills tests
SWE-Bench Verified
Accuracy • Self-reported
33.2%

Reasoning

Logical reasoning and analysis
GPQA
GPT-4o - Diamond no thinking no tools • Self-reported
70.1%

Multimodal

Working with images and visual data
MathVista
Accuracy • Self-reported
61.4%
MMMU
GPT-4o without mode thinking - Solution visual tasks level with • Self-reported
72.2%
ChartQA
evaluation on test set • Self-reported
85.7%
DocVQA
evaluation on test set • Self-reported
92.8%
AI2D
Evaluation on test set AI: Translate full text • Self-reported
94.2%

Other Tests

Specialized benchmarks
MMLU-Pro
0-shot CoT • Self-reported
74.7%
SimpleQA
accuracy • Self-reported
38.2%
AIME 2024
Accuracy AI • Self-reported
13.1%
IFEval
Accuracy AI • Self-reported
81.0%
Aider-Polyglot
Accuracy AI • Self-reported
30.7%
EgoSchema
evaluation on test set • Self-reported
72.2%
Aider-Polyglot Edit
Accuracy AI • Self-reported
18.2%
MMMLU
Accuracy • Self-reported
81.4%
Multi-IF
Accuracy • Self-reported
60.9%
TAU-bench Retail
Accuracy • Self-reported
60.3%
TAU-bench Airline
Accuracy • Self-reported
42.8%
CharXiv-R
GPT-4o without mode thinking - justification and • Self-reported
58.8%
Internal API instruction following (hard)
Accuracy • Self-reported
29.2%
MultiChallenge (o3-mini grader)
Accuracy AI: translation text by evaluation accuracy models AI: Accuracy • Self-reported
39.9%
COLLIE
GPT-4o without mode thinking - instructions at text • Self-reported
61.0%
Tau2 airline
GPT-4o without mode thinking - Benchmark functions () • Self-reported
45.5%
OpenAI-MRCR: 2 needle 128k
Accuracy AI • Self-reported
31.9%
Tau2 retail
GPT-4o without mode thinking - Benchmark functions () • Self-reported
63.4%
Tau2 telecom
GPT-4o without mode thinking - Benchmark functions (field) • Self-reported
23.5%
MMMU-Pro
GPT-4o without mode thinking - Solution visual tasks level with reasoning • Self-reported
59.9%
VideoMMMU
GPT-4o without mode thinking - multimodal reasoning (256 ) • Self-reported
61.2%
ERQA
GPT-4o without mode thinking - thinking • Self-reported
35.2%
Graphwalks parents <128k
Accuracy • Self-reported
35.4%
CharXiv-D
Accuracy AI: PageRank connection how therefore with number receive more rating When receives with she/it receives part "" this is "" to • Self-reported
85.3%
ComplexFuncBench
Accuracy • Self-reported
66.5%
SWE-Lancer
result • Self-reported
32.6%
SWE-Lancer (IC-Diamond subset)
score • Self-reported
12.4%
Graphwalks BFS <128k
Accuracy • Self-reported
41.7%
ActivityNet
evaluation on test set • Self-reported
61.9%
Humanity's Last Exam
GPT-4o without mode thinking (without tools) - set questions expert level by various subjects • Self-reported
5.3%
Scale MultiChallenge
GPT-4o without mode thinking - Benchmark execution instructions • Self-reported
40.3%
Multi-Challenge
GPT-4o without thinking mode - Multi-turn instruction following benchmark. • Self-reported
40.3%

License & Metadata

License
proprietary
Announcement Date
August 6, 2024
Last Updated
July 19, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.