BaiLian

Qwen3.6 Flash

AlibabaToken-based
Alias:qwen3.6-flash
Create API Key

Compare qwen3.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagevideofunction_callingtoolsstructured_outputjson_modeweb_searchcachingvisionreasoningcontext:1000000
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.169/ 1M
Output$1.014/ 1M
Cache Read$0.0169/ 1M
Alibaba
阿里巴巴百炼官方
Input$0.169/ 1M
Output$1.014/ 1M
Cache Read$0.0169/ 1M

Capabilities / Supported modalities

ReasoningFunction callingToolsStructured outputJSON modeWeb searchCode interpreter
Input
Output

Provider & data privacy

Provider
Alibaba (Qwen)Docs
Tokenizer
Qwen tokenizer (tiktoken-compat)
License
Tongyi Qianwen LicenseOpen weights
Data retention85 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3.6 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Qwen3.6 Flash

识别画面中的对象与位置关系

Qwen3.6 Flash 可理解图片和视频,也能处理文字、编程与推理任务。它加强了空间理解、物体定位和目标检测相关能力,适合从界面截图、商品陈列或操作画面中提取信息。例如识别按钮位于哪个区域、比较两个对象的相对位置,或根据连续画面整理操作顺序。

给视觉任务下指令时,应明确要找的主体与属性:问“右侧第三层有哪些红色包装商品”,比问“看到了什么”更便于得到有用结果。需要定位时,先要求描述可见特征与位置,再由应用检查目标是否正确;小字、遮挡和相似对象可通过补充局部清晰图来辅助判断。

从画面理解走向代码与工具

模型也适合编程智能体任务,可根据截图、代码和运行反馈提出修改,再借助应用提供的工具读取文件、执行检查或查询数据。比如定位弹窗遮挡问题时,同时提供组件代码、截图和预期交互,让模型把画面现象与布局逻辑联系起来。工具的实际执行结果应继续传回,作为下一步判断的依据。

它支持百万 Token 级上下文,可结合较长资料开展分析;仍建议保留任务相关内容,给不同材料加上编号。通过 enable_thinking 选择思考或直接回复:位置与逻辑联合判断可开启思考,简单标签提取可关闭。需要 JSON Mode 时使用非思考模式,并在程序中校验结果字段,便于接入内容标注、搜索和业务表单。

Use cases and prompting

例如分析网页截图,连同页面图片输入:

“请找到‘提交申请’按钮。描述按钮文字、颜色、所在区域,以及它与姓名输入框、取消按钮的相对位置。只依据当前截图;如果存在多个同名按钮,逐一说明区别。最后列出继续操作前仍缺少的信息,不猜测被遮挡的内容。”

需要结合组件代码排查点击失效时,再补充事件处理函数、控制台错误和复现步骤,并开启思考模式。仅做字段提取时可设置 enable_thinking=false,配合 JSON Mode 返回结构化结果。若应用据此执行点击,应先核对页面当前状态,操作后重新获取画面或工具结果确认变化。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about qwen3.6-flash

What is qwen3.6-flash?

Compare qwen3.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call qwen3.6-flash?

Create an API key with access to qwen3.6-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is qwen3.6-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of qwen3.6-flash?

The model catalog lists a context window of 1000000 tokens. Check the selected endpoint for request limits.

What is the maximum output of qwen3.6-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate qwen3.6-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.