Google logo

Gemini 3.1 Pro

Multimodal
Google

Gemini 3.1 Pro is the latest model in the Gemini 3 series. Excels at complex tasks requiring extensive world knowledge and advanced multimodal reasoning. Gemini 3.1 Pro uses dynamic thinking by default to process prompts and has a 1 million token input context window with 64k output tokens.

Key Specifications

Parameters
-
Context
1.0M
Release Date
February 19, 2026
Average Score
63.3%

Timeline

Key dates in the model's history
Announcement
February 19, 2026
Last Update
February 20, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
January 1, 2025
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.50
Output (per 1M tokens)
$15.00
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-bench Verified
SWE-bench Verified — benchmark for evaluation abilities model solve real tasks from GitHub- • Self-reported
80.6%

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond — benchmark for evaluation abilities model answer on questions level PhD by and • Self-reported
94.3%

Other Tests

Specialized benchmarks
ARC-AGI v2
ARC-AGI v2 — benchmark for evaluation abilities to reasoning and generalization • Self-reported
77.1%
MMMLU
MMMLU — version MMLU for evaluation knowledge model on languages • Self-reported
92.6%
CharXiv-R
CharXiv-R — benchmark for evaluation abilities model understand and reason about and • Self-reported
85.9%
MMMU-Pro
MMMU-Pro — version MMMU for evaluation on level experts • Self-reported
80.5%
HLE
HLE (Humanity's Last Exam) — benchmark from questions, experts for verification knowledge AI • Self-reported
51.4%
APEX-Agents
• Self-reported
33.5%
BrowseComp
Search + Python + Browse • Self-reported
85.9%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, high effort; Pass@1 12% ± 2%. • Verified
12.0%
Finance Agent v2
• Verified
43.0%
FrontierSWE
Gemini CLI • Verified
40.0%
Legal Agent Benchmark
All-pass rate (Harvey held-out set) • Verified
-
LiveBench
2026-01-08, High • Verified
79.9%
LiveCodeBench Pro
Elo Rating • Self-reported
96.2%
MCP Atlas
• Self-reported
69.2%
MRCR v2 (8-needle)
1M (pointwise) • Self-reported
26.3%
SciCode
• Self-reported
59.0%
SWE-Bench Pro
Single attempt • Self-reported
54.2%
t2-bench
Telecom • Self-reported
99.3%
Terminal-Bench 2.0
Terminus-2 harness • Self-reported
68.5%

License & Metadata

License
proprietary
Announcement Date
February 19, 2026
Last Updated
February 20, 2026

Articles about Gemini 3.1 Pro

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.