Zhipu AI logo

GLM-5V-Turbo

Multimodal
Zhipu AI

GLM-5V-Turbo is Z.AI's first multimodal coding foundation model, built for vision-based coding tasks. It natively processes multimodal inputs including images, video, text, and files, while excelling at long-horizon planning, complex coding, and action execution.

Key Specifications

Parameters
-
Context
200.0K
Release Date
April 2, 2026
Average Score
65.0%

Timeline

Key dates in the model's history
Announcement
April 2, 2026
Last Update
August 29, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.20
Output (per 1M tokens)
$4.00
Max Input Tokens
200.0K
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
AndroidWorld
Self-reported
75.7%
BrowseComp-VL
Self-reported
51.9%
CC-Bench-V2 Backend
Self-reported
22.8%
CC-Bench-V2 Frontend
Self-reported
68.4%
CC-Bench-V2 Repo Exploration
Self-reported
72.2%
Claw-Eval
Pass@3Self-reported
75.0%
Design2Code
Self-reported
94.8%
FACTS Grounding
Self-reported
58.6%
Flame-VLM-Code
Self-reported
93.8%
ImageMining
Self-reported
30.7%
MMSearch
Self-reported
72.9%
MMSearch-Plus
Self-reported
30.0%
OSWorld
Self-reported
62.3%
PinchBench
AverageSelf-reported
80.7%
SimpleVQA
Self-reported
78.2%
V*
Self-reported
89.0%
Vision2Web
Self-reported
31.0%
WebVoyager
Self-reported
88.5%
ZClawBench
Self-reported
57.6%

License & Metadata

License
proprietary
Announcement Date
April 2, 2026
Last Updated
August 29, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.