Z.ai

GLM 5.3 Flash

Z.aiToken-based
Alias:glm-5.3-flash
Create API Key

Compare glm-5.3-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagevideofilestreamingfunction_callingtoolsjson_modevisionreasoningcachingcontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
943.7K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.12/ 1M
Output$0.418/ 1M
Cache Read$0.035/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
Zhipu AIDocs
Tokenizer
GLM tokenizer
License
GLM-4 LicenseOpen weights
Data retention60 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·GLM 5.3 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About GLM 5.3 Flash

把视觉材料带进工程工作

GLM-5.3-Flash 在GLM-5系列中采用原生多模态设计,可以结合文字、图片和视频理解任务,输出文本结果。它适合需要看懂界面、图表或操作过程的编程与智能体应用,也适合将多种素材整理成可执行的工作说明。

适合边看边分析的场景

例如维护一个管理系统,可以把页面截图与相关代码一起提供,让模型解释界面状态和实现之间的关系;整理操作流程时,可以结合录像与文字手册,找出遗漏步骤或容易出错的环节。视觉素材应围绕明确问题提交,说明要关注的对象、区域或动作,避免只得到宽泛的画面描述。

模型面向能力与效率兼顾的工作,可用于多轮编程、资料整理与自动化流程。应用接入工具后,可以把分析连接到实际查询、执行和检查;每次工具结果都应进入后续上下文,帮助模型决定下一步。

思考强度需要明确选择

GLM-5.3-Flash支持low、high和max推理等级,未指定时默认max。想减少思考投入时,应明确传入受支持的等级,再用真实任务比较效果。长材料任务要保留关键约束与出处,复杂视觉判断则应检查原图或录像,避免把模糊区域当作确定事实。

Use cases and prompting

操作手册检查示例

“结合这段后台操作录像、页面截图和现有手册,检查创建项目流程。列出录像实际经过的步骤,指出手册遗漏或顺序不同的地方,再写一版简洁操作说明。看不清的控件名称标为待确认,不要自行编造。”

代码任务可同时提供相关组件,使问题能对应到实现。

想加快简单任务,为什么没明显变化?

先确认明确传入low或high。该模型未指定推理等级时默认max,其他不受支持的值也不能代替正确设置。

图片和视频应该怎么搭配?

视频用于理解过程,清晰截图用于补充关键细节;在提示词中说明两者的关系和关注位置。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about glm-5.3-flash

What is glm-5.3-flash?

Compare glm-5.3-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call glm-5.3-flash?

Create an API key with access to glm-5.3-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is glm-5.3-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of glm-5.3-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of glm-5.3-flash?

The model catalog lists a maximum output of 943717 tokens. Your request settings may set a lower limit.

How should I evaluate glm-5.3-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.