Key Specifications
Parameters
-
Context
16.4K
Release Date
March 21, 2023
Average Score
42.3%
Timeline
Key dates in the model's history
Announcement
March 21, 2023
Last Update
July 19, 2025
Today
October 5, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
September 30, 2021
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$0.50
Output (per 1M tokens)
$1.50
Max Input Tokens
16.4K
Max Output Tokens
4.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
Accuracy
AI • Verified
Programming
Programming skills tests
HumanEval
Accuracy • Verified
Mathematics
Mathematical problems and computations
MATH
Accuracy • Verified
MGSM
Accuracy
AI: Human • Verified
Reasoning
Logical reasoning and analysis
GPQA
Accuracy • Verified
DROP
Accuracy • Verified
Multimodal
Working with images and visual data
MathVista
Accuracy AI: still but I how Stability AI and Anthropic (in ) make large steps Models level Gorilla have accuracy use API, than and Anthropic that Claude can more exactly perform instructions. I that accuracy answers • Verified
MMMU
Accuracy • Verified
License & Metadata
License
proprietary
Announcement Date
March 21, 2023
Last Updated
July 19, 2025
Similar Models
All Modelso3-mini
OpenAI
Best score:0.9 (MMLU)
Released:Jan 2025
Price:$1.10/1M tokens
GPT-5 Codex
OpenAI
Released:Sep 2025
Price:$2.00/1M tokens
GPT-4 Turbo
OpenAI
Best score:0.9 (HumanEval)
Released:Apr 2024
Price:$10.00/1M tokens
o1-mini
OpenAI
Best score:0.9 (HumanEval)
Released:Sep 2024
Price:$3.00/1M tokens
o1
OpenAI
Best score:0.9 (MMLU)
Released:Dec 2024
Price:$15.00/1M tokens
o1-preview
OpenAI
Best score:0.9 (MMLU)
Released:Sep 2024
Price:$15.00/1M tokens
Claude 3.5 Haiku
Anthropic
Best score:0.9 (HumanEval)
Released:Oct 2024
Price:$0.80/1M tokens
Qwen3 Max
Alibaba
Best score:0.6 (GPQA)
Released:Dec 2025
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.