Google logo

Gemini 3.5 Flash

Multimodal
Google

Gemini 3.5 Flash is Google's fast frontier model in the Gemini 3.5 series, optimized for low latency and efficient scaling across chat, agentic, and multimodal workloads. It delivers strong performance on tool use and multimodal reasoning benchmarks such as MCP Atlas and MMMU-Pro. The model supports a 1M token input context window with up to 65.5K output tokens.

Key Specifications

Parameters
-
Context
1.0M
Release Date
May 19, 2026
Average Score
57.4%

Timeline

Key dates in the model's history
Announcement
May 19, 2026
Last Update
August 27, 2026
Today
September 8, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
January 31, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.50
Output (per 1M tokens)
$9.00
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
CharXiv-R
No toolsSelf-reported
84.0%
MCP Atlas
Self-reported
84.0%
MMMU-Pro
No toolsSelf-reported
84.0%
OSWorld-Verified
Self-reported
78.0%
ARC-AGI v2
Self-reported
72.1%
Blueprint-Bench 2
Normalized scoreSelf-reported
33.6%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, medium effort; Pass@1 37% ± 2%.Verified
37.0%
Finance Agent
Finance Agent v2Self-reported
57.9%
Finance Agent v2
Verified
57.9%
Humanity's Last Exam
Self-reported
40.2%
Legal Agent Benchmark
All-pass rate (Harvey held-out set)Verified
0.8%
LiveBench
2026-01-08, HighVerified
75.0%
MRCR v2 (8-needle)
1M (pointwise)Self-reported
26.6%
SWE-Bench Pro
SWE-Bench Pro (Public). Single attemptSelf-reported
55.1%
Terminal-Bench 2.0
Terminus-2 harness (Terminal-Bench 2.1)Self-reported
76.2%
Toolathlon
Self-reported
56.5%

License & Metadata

License
proprietary
Announcement Date
May 19, 2026
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.