Z.ai

GLM 5.3

Z.aiToken-based
Alias:glm-5.3
Create API Key

Compare glm-5.3 API pricing, supported endpoints, capabilities and access options on Modelsell.

textstreamingfunction_callingtoolsjson_modereasoningcachingcontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
943.7K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$1.2/ 1M
Output$4.18/ 1M
Cache Read$0.3/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoning
Input
Output

Provider & data privacy

Provider
Zhipu AIDocs
Tokenizer
GLM tokenizer
License
GLM-4 LicenseOpen weights
Data retention89 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·GLM 5.3

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About GLM 5.3

处理一段代码之外的工程问题

GLM-5.3 是面向复杂软件工程和长周期智能体任务的推理模型。它适合阅读项目结构后制定改造方案,分析模块间依赖,处理代码迁移,或围绕既定目标持续完成实现与验证。对于需要把很多细节同时放在心里的工作,它可以帮助组织分析顺序和实施步骤。

长材料要有清楚的任务主线

模型支持文本输入和输出,可结合代码、接口说明、变更记录与错误日志进行分析。代码迁移时,先说明要保留的行为和兼容范围,再提供目标环境与验收条件;问题排查时,则保留调用路径和失败案例,让结论能对应具体证据。

跨模块工作适合按依赖顺序推进。先识别影响面,再完成最小可验证改动,检查通过后继续下一部分。应用接入执行工具时,应把编译结果、测试失败和运行状态交回,帮助模型修正后续计划。没有执行条件时,也应区分建议方案与已验证结果。

始终开启的推理

GLM-5.3的推理始终开启,可用low、high和max调整投入。简单任务可以减少分析深度,复杂工程任务则需要留出足够验证空间。长流程中保存关键决策、已完成修改与剩余风险,可以减少重复工作,也便于团队在中途审阅方向。

Use cases and prompting

代码迁移示例

“将现有文件存储模块迁移到新的客户端库。先找出所有调用方和行为差异,保留上传、下载、重试及错误返回约定。按依赖顺序分步改动,每步验证成功路径、超时和权限失败,最后列出仍依赖旧库的位置。”

提供库版本、相关代码与兼容要求,避免只给一个笼统迁移目标。

可以完全关闭推理吗?

GLM-5.3始终启用推理,可选择low、high或max调整强度。

多模块改造怎样避免遗漏?

先建立调用清单和验收项,分阶段记录已修改位置与验证结果。最终对照清单检查,不能只以代码生成完成作为结束标志。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about glm-5.3

What is glm-5.3?

Compare glm-5.3 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call glm-5.3?

Create an API key with access to glm-5.3, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is glm-5.3 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of glm-5.3?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of glm-5.3?

The model catalog lists a maximum output of 943718 tokens. Your request settings may set a lower limit.

How should I evaluate glm-5.3 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.