Zhipu AI logo

GLM-5.3

Zhipu AI

GLM-5.3 is Zhipu AI's 753B-parameter open-weight MoE model released in August 2026, with a 1M token context window. It is tuned for agentic workflows, showing strong results on Terminal-Bench 2.1, CyberGym and FrontierSWE. It is served by Novita and Z.ai.

Key Specifications

Parameters
753.0B
Context
1.0M
Release Date
August 14, 2026
Average Score
52.2%

Timeline

Key dates in the model's history
Announcement
August 14, 2026
Last Update
August 27, 2026
Today
September 8, 2026

Technical Specifications

Parameters
753.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$1.40
Output (per 1M tokens)
$4.40
Max Input Tokens
1.0M
Max Output Tokens
131.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
Terminal-Bench 2.1
Claude Code 2.1.207; max reasoning effortSelf-reported
88.0%
CyberGym
Claude Code 2.1.207; max reasoning effort; single-run Pass@1 over 1,507 tasksSelf-reported
84.5%
FrontierSWE
Claude Code 2.1.207; max reasoning effortSelf-reported
78.0%
Agents' Last Exam
ALE-CLI; Claude Code harness; max reasoning effort; 1M-token contextSelf-reported
28.5%
AutomationBench
AutomationBench v1.0.6Self-reported
48.2%
DeepSWE 1.1
DeepSWE v1.1; mini-swe-agent harness; max reasoning effortSelf-reported
66.9%
ExploitBench
Claude Code 2.1.207; max reasoning effort; average coverage over 41 tasks and 3 revisionsSelf-reported
54.4%
ExploitGym
Six-hour normalized inference-time budget; Claude Code 2.1.207; max reasoning effortSelf-reported
15.0%
GDPval-AA
GDPval-AA v2 Elo (max_score 3000)Self-reported
59.0%
Humanity's Last Exam
With tools; 300K-token context; GPT-5.6 Luna judgeSelf-reported
62.5%
NL2Repo
1M-token context; rule-based and LLM-based anti-hacking checksSelf-reported
58.0%
PostTrainBench
Claude Code 2.1.207; max reasoning effort; weighted average of 3 runsSelf-reported
39.8%
Program Bench
Almost Solved metricSelf-reported
19.0%
SWE-Marathon
SWE-Marathon v1.1; Claude Code 2.1.207; max reasoning effortSelf-reported
42.5%
Terminal-Bench 3.0
Claude Code 2.1.207; max reasoning effortSelf-reported
28.3%
Terminal-Bench 4.0
Terminal-Bench 4.0 public leaderboard; resolution rate 41.8% ± 3.2%; Claude Code agent; max effort.Verified
41.8%
Toolathlon
Toolathlon VerifiedSelf-reported
73.0%

License & Metadata

License
mit
Announcement Date
August 14, 2026
Last Updated
August 27, 2026

Compare GLM-5.3

All comparisons

Articles about GLM-5.3

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.