Z.ai

GLM 5

Z.aiToken-based
Alias:glm-5
Create API Key

Compare glm-5 API pricing, supported endpoints, capabilities and access options on Modelsell.

textfunction_callingtoolsjson_modereasoningcontext:204800
Starting price
Input / Output · 1M
Context
204.8K
Maximum input window
Max output
128K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.58/ 1M
Output$2.6/ 1M
Cache Read$0.14/ 1M
Alibaba
阿里巴巴百炼官方
Input$0.58/ 1M
Output$2.6/ 1M
Cache Read$0.14/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoning
Input
Output

Provider & data privacy

Provider
Zhipu AIDocs
Tokenizer
GLM tokenizer
License
GLM-4 LicenseOpen weights
Data retention67 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·GLM 5

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About GLM 5

先理解系统,再解决局部问题

GLM-5关注复杂系统工程和长周期智能体任务,适合需要同时考虑架构、接口、状态与异常处理的工作。它可以帮助研发团队梳理模块关系、分析故障路径、规划功能实现,也可以围绕一组资料完成持续的推理与信息整合。

当问题横跨多个组件时,单看报错位置往往不够。可以把请求入口、业务逻辑、数据访问和任务处理的相关实现放在一起,让模型沿实际流程判断原因。描述目标时,应说明正常行为、失败现象与必须保留的约束,这能让方案更贴近已有系统。

适合有反馈的工作方式

GLM-5可以参与代码生成和工具协作,但有效的工程流程仍需要真实反馈。应用执行查询、编译或测试后,应把输出交回模型,保留失败信息,让下一轮改动围绕实际问题展开。先做一个最小验证,再逐步扩大改动范围,有助于减少无关重构。

模型以文本输入和输出为主,适合代码、日志、文档和结构化信息。面对长任务,可以记录已确认事实、实施步骤和待验证假设;完成时要求说明实际产物与验证结果,方便团队继续审阅和维护。

Use cases and prompting

异步任务设计示例

“为图片处理服务设计异步任务流程。要求提交后返回任务标识,处理中可查询状态,失败可安全重试。结合现有代码说明状态如何流转,怎样避免重复处理,哪些错误可以重试,并给出最小实施顺序和验证用例。”

把现有接口、队列行为与数据约束一起提供,避免得到脱离项目的通用方案。

生成的方案太大,怎样收敛?

明确本次只解决的问题、允许修改的模块和交付期限,要求先完成最小可用流程。

怎样检查它有没有真正理解系统?

先让它复述一次完整请求路径,并指出关键状态和失败点,再开始改动;发现理解偏差时及时补充材料。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about glm-5

What is glm-5?

Compare glm-5 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call glm-5?

Create an API key with access to glm-5, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is glm-5 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of glm-5?

The model catalog lists a context window of 204800 tokens. Check the selected endpoint for request limits.

What is the maximum output of glm-5?

The model catalog lists a maximum output of 128000 tokens. Your request settings may set a lower limit.

How should I evaluate glm-5 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.