Google logo

Gemma 4 12B

Multimodal
Google

Gemma 4 12B is Google DeepMind's encoder-free multimodal instruction-tuned model with 11.95 billion parameters and a 256K context window. It supports text, image, audio, and video inputs with text output, projecting image patches and audio waveforms directly into a single decoder-only transformer for streamlined local deployment.

Key Specifications

Parameters
12.0B
Context
-
Release Date
May 23, 2026
Average Score
59.4%

Timeline

Key dates in the model's history
Announcement
May 23, 2026
Last Update
September 10, 2026
Today
September 19, 2026

Technical Specifications

Parameters
12.0B
Training Tokens
-
Knowledge Cutoff
January 1, 2025
Family
-
Capabilities
MultimodalZeroEval

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
Self-reported
78.8%

Other Tests

Specialized benchmarks
AIME 2026
No toolsSelf-reported
77.5%
BIG-Bench Extra Hard
Self-reported
53.0%
CodeForces
Elo ratingSelf-reported
55.3%
CoVoST2
BLEUSelf-reported
38.5%
FLEURS
Speech recognition accuracy (1 - WER)Self-reported
93.1%
Humanity's Last Exam
No toolsSelf-reported
5.2%
LiveCodeBench v6
Self-reported
72.0%
MathVision
Self-reported
79.7%
MedXpertQA
MMSelf-reported
48.7%
MMLU-Pro
Self-reported
77.2%
MMMLU
Self-reported
83.4%
MMMU-Pro
Self-reported
69.1%
MRCR v2 (8-needle)
128kSelf-reported
43.4%
OmniDocBench 1.5
Average edit distanceSelf-reported
16.4%

License & Metadata

License
apache_2_0
Announcement Date
May 23, 2026
Last Updated
September 10, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.