OpenAI logo

GPT-4.1 mini

Multimodal
OpenAI

GPT-4.1 mini balances intelligence, speed, and cost. It represents a significant breakthrough in small model performance, even surpassing GPT-4o on many benchmarks while reducing latency and cost.

Key Specifications

Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
49.6%

Timeline

Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
May 31, 2024
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.40
Output (per 1M tokens)
$1.60
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Standard benchmark AI: whether more complex this tasks? • Self-reported
87.5%

Programming

Programming skills tests
SWE-Bench Verified
methodology, on [2] • Self-reported
23.6%

Reasoning

Logical reasoning and analysis
GPQA
Diamond • Self-reported
65.0%

Multimodal

Working with images and visual data
MathVista
Standard benchmark AI: Human_Evaluation • Self-reported
73.1%
MMMU
Standard benchmark • Self-reported
72.7%

Other Tests

Specialized benchmarks
AIME 2024
Standard benchmark AI: I'll solve this step-by-step. • Self-reported
49.6%
IFEval
Standard benchmark • Self-reported
84.1%
Aider-Polyglot
Standard benchmark • Self-reported
34.7%
MultiChallenge
Standard benchmark (GPT-4o grader) • Self-reported
35.8%
Aider-Polyglot Edit
Standard benchmark AI: (1/8) Standard benchmark • Self-reported
31.6%
MMMLU
Standard benchmark AI: (7 of 25 marks) • Self-reported
78.5%
Multi-IF
Standard benchmark AI: I'm ChatGPT, a large language model trained by OpenAI, based on the GPT-4 architecture. • Self-reported
67.0%
TAU-bench Retail
Average value by 5 without tools/prompts (note [4], model GPT-4o) • Self-reported
55.8%
TAU-bench Airline
Average by 5 without use special tools/prompts ([4]) • Self-reported
36.0%
CharXiv-R
Standard benchmark • Self-reported
56.8%
Internal API instruction following (hard)
Internal benchmark AI: *internal thoughts* This too text for translation. I its exactly Internal benchmark • Self-reported
45.1%
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini, [3]) • Self-reported
42.2%
COLLIE
Standard benchmark AI: (thinking) Standard benchmark = standard benchmark. "benchmark" without translation, so how this in AI • Self-reported
54.6%
OpenAI-MRCR: 2 needle 128k
Internal benchmark • Self-reported
47.2%
OpenAI-MRCR: 2 needle 1M
Internal benchmark • Self-reported
33.3%
Graphwalks BFS <128k
Standard benchmark Standard benchmark AI • Self-reported
61.7%
Graphwalks BFS >128k
Internal benchmark • Self-reported
15.0%
Graphwalks parents <128k
Internal benchmark • Self-reported
60.5%
Graphwalks parents >128k
Internal benchmark • Self-reported
11.0%
CharXiv-D
Standard benchmark AI: on this task. First let's solve her/its method. Task: $x$ such, that $2^x = 32$. In order to find $x$, I I can use $2^x = 32$ $2^x = 2^5$ (so how $32 = 2^5$) should scores, therefore $x = 5$. Answer: $x = 5$ • Self-reported
88.4%
ComplexFuncBench
Standard benchmark • Self-reported
49.3%
AIME 2025
GPT-4.1 mini without tools - mathematics (AIME 2025) • Self-reported
40.2%
Humanity's Last Exam
GPT-4.1 mini without tools - Questions expert level by various subjects. • Self-reported
3.7%
HMMT 2025
GPT-4.1 mini without tools - Harvard-MIT Mathematics Tournament. • Self-reported
35.0%

License & Metadata

License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.