Google logo

Gemini 3.5 Flash

Multimodal
Google

Gemini 3.5 Flash is Google's fast frontier model in the Gemini 3.5 series, optimized for low latency and efficient scaling across chat, agentic, and multimodal workloads. It delivers strong performance on tool use and multimodal reasoning benchmarks such as MCP Atlas and MMMU-Pro. The model supports a 1M token input context window with up to 65.5K output tokens.

Key Specifications

Parameters
-
Context
1.0M
Release Date
May 19, 2026
Average Score
82.5%

Timeline

Key dates in the model's history
Announcement
May 19, 2026
Last Update
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
January 31, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.50
Output (per 1M tokens)
$9.00
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
CharXiv-R
No toolsSelf-reported
84.0%
MCP Atlas
Self-reported
84.0%
MMMU-Pro
No toolsSelf-reported
84.0%
OSWorld-Verified
Self-reported
78.0%

License & Metadata

License
proprietary
Announcement Date
May 19, 2026
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.