BaiLian

Qwen3.5 Flash

AlibabaToken-based
Alias:qwen3.5-flash
Create API Key

Compare qwen3.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagevideofunction_callingtoolsstructured_outputjson_modeweb_searchcachingvisionreasoningcontext:1000000
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.0282/ 1M
Output$0.282/ 1M
Cache Read$0.00282/ 1M
Alibaba
阿里巴巴百炼官方
Input$0.0282/ 1M
Output$0.282/ 1M
Cache Read$0.00282/ 1M

Capabilities / Supported modalities

ReasoningFunction callingToolsStructured outputJSON modeWeb searchCode interpreter
Input
Output

Provider & data privacy

Provider
Alibaba (Qwen)Docs
Tokenizer
Qwen tokenizer (tiktoken-compat)
License
Tongyi Qianwen LicenseOpen weights
Data retention70 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3.5 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Qwen3.5 Flash

把文字与画面整理成可用信息

Qwen3.5 Flash 适合高频的内容处理与交互任务,例如把用户反馈整理为工单、根据界面截图解释操作问题、概括视频中发生的事情,或从多份资料中提取统一字段。它接受文字、图片和视频,输出文字,也能按指定结构组织结果,便于后续展示、筛选和程序处理。

在产品支持场景中,可以同时提供用户描述、报错截图和帮助文档,让模型区分画面上能看到的事实、用户陈述与尚待确认的信息。整理长资料时,给每份材料设置编号,并要求结论带上对应编号,有助于回到原文检查。它支持较长的上下文,但清楚的任务目标和相关材料仍比堆入无关内容更有用。

哪些任务需要开启思考?

通过 enable_thinking 可切换思考模式。简单分类、格式整理和字段提取,可以关闭思考;需要比较多个方案、推断故障原因或分析前后矛盾时,可以开启。思考会增加处理过程与输出用量,应用应为最终答案留足预算,并在流式处理中区分思考内容与正式回复。

如何接入业务流程?

模型支持函数调用,适合根据对话选择查询订单、读取知识库等工具。应用执行工具并传回结果后,模型再组织答复;联网搜索等内置工具也需要在请求中启用。要求 JSON 时,除说明字段含义外,还应使用可用的结构化输出设置,并在程序中检查字段、类型与必填信息。

Use cases and prompting

例如整理界面故障,传入截图与用户描述后输入:

“请生成一条产品工单。先列出截图中可直接观察到的问题,再概括用户诉求;推测原因另列,不把猜测写成事实。输出包含 title、observations、user_request、missing_info 四个字段,observations 和 missing_info 为数组。看不清的文字写‘无法辨认’。用户描述:点击导出后一直停在加载中。”

常规字段提取可设置 enable_thinking=false;需要结合日志与多个页面定位原因时设置为 true。使用 Python OpenAI SDK 时,通过 extra_body 传递该扩展参数。图片需作为实际图片内容随消息提交,程序收到结果后再校验 JSON 并保存工单。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about qwen3.5-flash

What is qwen3.5-flash?

Compare qwen3.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call qwen3.5-flash?

Create an API key with access to qwen3.5-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is qwen3.5-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of qwen3.5-flash?

The model catalog lists a context window of 1000000 tokens. Check the selected endpoint for request limits.

How should I evaluate qwen3.5-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.