OpenAI logo

GPT-4.1 mini

Multimodal
OpenAI

GPT-4.1 mini balances intelligence, speed, and cost. It represents a significant breakthrough in small model performance, even surpassing GPT-4o on many benchmarks while reducing latency and cost.

Key Specifications

Parameters
-
Context
1.0M
Release Date
April 14, 2025
Average Score
49.6%

Timeline

Key dates in the model's history
Announcement
April 14, 2025
Last Update
July 19, 2025
Today
September 10, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
May 31, 2024
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.40
Output (per 1M tokens)
$1.60
Max Input Tokens
1.0M
Max Output Tokens
32.8K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
Standard benchmark AI: whether more complex this tasks?Self-reported
87.5%

Programming

Programming skills tests
SWE-Bench Verified
methodology, on [2]Self-reported
23.6%

Reasoning

Logical reasoning and analysis
GPQA
DiamondSelf-reported
65.0%

Multimodal

Working with images and visual data
MathVista
Standard benchmark AI: Human_EvaluationSelf-reported
73.1%
MMMU
Standard benchmarkSelf-reported
72.7%

Other Tests

Specialized benchmarks
AIME 2024
Standard benchmark AI: I'll solve this step-by-step.Self-reported
49.6%
IFEval
Standard benchmarkSelf-reported
84.1%
Aider-Polyglot
Standard benchmarkSelf-reported
34.7%
MultiChallenge
Standard benchmark (GPT-4o grader)Self-reported
35.8%
Aider-Polyglot Edit
Standard benchmark AI: (1/8) Standard benchmarkSelf-reported
31.6%
MMMLU
Standard benchmark AI: (7 of 25 marks)Self-reported
78.5%
Multi-IF
Standard benchmark AI: I'm ChatGPT, a large language model trained by OpenAI, based on the GPT-4 architecture.Self-reported
67.0%
TAU-bench Retail
Average value by 5 without tools/prompts (note [4], model GPT-4o)Self-reported
55.8%
TAU-bench Airline
Average by 5 without use special tools/prompts ([4])Self-reported
36.0%
CharXiv-R
Standard benchmarkSelf-reported
56.8%
Internal API instruction following (hard)
Internal benchmark AI: *internal thoughts* This too text for translation. I its exactly Internal benchmarkSelf-reported
45.1%
MultiChallenge (o3-mini grader)
Standard benchmark (o3-mini, [3])Self-reported
42.2%
COLLIE
Standard benchmark AI: (thinking) Standard benchmark = standard benchmark. "benchmark" without translation, so how this in AISelf-reported
54.6%
OpenAI-MRCR: 2 needle 128k
Internal benchmarkSelf-reported
47.2%
OpenAI-MRCR: 2 needle 1M
Internal benchmarkSelf-reported
33.3%
Graphwalks BFS <128k
Standard benchmark Standard benchmark AISelf-reported
61.7%
Graphwalks BFS >128k
Internal benchmarkSelf-reported
15.0%
Graphwalks parents <128k
Internal benchmarkSelf-reported
60.5%
Graphwalks parents >128k
Internal benchmarkSelf-reported
11.0%
CharXiv-D
Standard benchmark AI: on this task. First let's solve her/its method. Task: $x$ such, that $2^x = 32$. In order to find $x$, I I can use $2^x = 32$ $2^x = 2^5$ (so how $32 = 2^5$) should scores, therefore $x = 5$. Answer: $x = 5$Self-reported
88.4%
ComplexFuncBench
Standard benchmarkSelf-reported
49.3%
AIME 2025
GPT-4.1 mini without tools - mathematics (AIME 2025)Self-reported
40.2%
Humanity's Last Exam
GPT-4.1 mini without tools - Questions expert level by various subjects.Self-reported
3.7%
HMMT 2025
GPT-4.1 mini without tools - Harvard-MIT Mathematics Tournament.Self-reported
35.0%

License & Metadata

License
proprietary
Announcement Date
April 14, 2025
Last Updated
July 19, 2025

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.