Google logo

Gemini 3.1 Flash-Lite

Multimodal
Google

Gemini 3.1 Flash-Lite is the first Flash-Lite model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, content moderation, and classification. It delivers strong performance at a fraction of the cost of larger Gemini models, with 2.5x faster Time to First Token compared to the previous generation. Supports a 1 million token input context with 64k output tokens.

Key Specifications

Parameters
-
Context
1.0M
Release Date
March 3, 2026
Average Score
54.6%

Timeline

Key dates in the model's history
Announcement
March 3, 2026
Last Update
September 1, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
January 1, 2025
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.25
Output (per 1M tokens)
$1.50
Max Input Tokens
1.0M
Max Output Tokens
65.5K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Reasoning

Logical reasoning and analysis
GPQA
No toolsSelf-reported
86.9%

Other Tests

Specialized benchmarks
CharXiv-R
ScoreSelf-reported
73.2%
FACTS Grounding
ScoreSelf-reported
40.6%
Finance Agent v2
ScoreVerified
30.0%
Humanity's Last Exam
No toolsSelf-reported
16.0%
Legal Agent Benchmark
All-pass rate (Harvey held-out set)Verified
-
MMMLU
ScoreSelf-reported
88.9%
MMMU-Pro
No toolsSelf-reported
76.8%
MRCR v2 (8-needle)
128k (average)Self-reported
60.1%
SimpleQA
ScoreSelf-reported
43.3%
VideoMMMU
ScoreSelf-reported
84.8%

License & Metadata

License
proprietary
Announcement Date
March 3, 2026
Last Updated
September 1, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.