BaiLian

Qwen3 Omni Flash 2025 12.01

AlibabaToken-based
Alias:qwen3-omni-flash-2025-12-01
Create API Key

Compare qwen3-omni-flash-2025-12-01 API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideo
Starting price
Input / Output · 1M
Context
—
Maximum input window
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$75/ 1M
Output$75/ 1M
Alibaba
阿里巴巴百炼官方
Input$75/ 1M
Output$75/ 1M

Capabilities / Supported modalities

StreamingSystem promptFunction callingToolsJSON modeStructured outputVision
Input
Output

Provider & data privacy

Provider
Alibaba (Qwen)Docs
Tokenizer
Qwen tokenizer (tiktoken-compat)
License
Tongyi Qianwen LicenseOpen weights
Data retention52 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3 Omni Flash 2025 12.01

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Qwen3 Omni Flash 2025 12.01

听懂内容,也能用语音回答

Qwen3 Omni Flash 可以理解文字、图片、音频和视频,并生成文字或自然语音答复。它适合分析一段演示视频、整理录音中的主要信息,或根据图片回答问题。也可以通过系统提示设定角色、语气和回答长度,让多轮交互保持清楚的表达风格。

例如上传一段产品演示视频,让它说明主要操作与讲解内容,再整理成简短使用说明;上传一段英语语音,让它用中文概括要点,或用支持的语音语言回答。这里的语音输出是在理解内容后生成回复,不是给原录音做声音复刻。

输入材料怎样组合?

这一代模型支持文字与一种其他模态组合输入,如文字加图片、文字加音频或文字加视频。音视频应控制在 150 秒以内。长内容可以先切成有意义的片段,并在文字中说明前后关系;不要把多个独立媒体类型随意混成同一轮输入。

文字分析与语音回复怎么选?

模型支持通过 enable_thinking 开启思考,适合需要进一步分析的问题;思考模式只输出文字。需要语音回复时关闭思考,选择文字加音频输出,并设置音色。调用使用流式返回,应用需要分别接收文字与音频数据,再播放生成的声音。

它可以在支持函数调用的接入中连接应用工具,但没有内置联网搜索。需要实时资料时,应由应用获取并传入。处理听不清的语音或模糊画面时,让模型指出不确定内容,避免将猜测写进转录、摘要或操作说明。

Use cases and prompting

分析演示视频时,先准备 150 秒以内的片段,再输入:

“请按操作顺序概括视频中的步骤,区分画面可见操作与讲解者说的话;没有展示结果的步骤不要写成成功。最后生成一段适合新人听的简短说明。”

只要文字分析时,将 modalities 设置为仅包含 text;需要深入分析,可开启 enable_thinking。要同时获得语音答复,关闭思考,将 modalities 设置为包含 text 和 audio,在 audio 中指定支持的音色,例如 Cherry,以及音频格式,例如 wav,并启用 stream: true。程序分别收集文字和音频流,检查内容后播放。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about qwen3-omni-flash-2025-12-01

What is qwen3-omni-flash-2025-12-01?

Compare qwen3-omni-flash-2025-12-01 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call qwen3-omni-flash-2025-12-01?

Create an API key with access to qwen3-omni-flash-2025-12-01, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is qwen3-omni-flash-2025-12-01 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

How should I evaluate qwen3-omni-flash-2025-12-01 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.