BaiLian

Qwen3 Omni Flash Realtime

AlibabaToken-based
Alias:qwen3-omni-flash-realtime
Create API Key

Compare qwen3-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.

audiotextimagevideorealtimestreaming多模态对话
Starting price
Input / Output · 1M
Context
65.5K
Maximum input window
Max output
16.4K
Maximum tokens per response
Modalities
→
Released
Dec 2025

Pricing by Supplier

official
官方接口直连
Input$3.584/ 1M
Output$7.168/ 1M
Alibaba
阿里巴巴百炼官方
Input$3.584/ 1M
Output$7.168/ 1M

Capabilities / Supported modalities

StreamingVisionSystem prompt
Input
Output

Provider & data privacy

Provider
Alibaba (Qwen)Docs
Tokenizer
Qwen tokenizer (tiktoken-compat)
License
Tongyi Qianwen LicenseOpen weights
Data retention57 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3 Omni Flash Realtime

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Qwen3 Omni Flash Realtime

一边交流,一边理解画面

Qwen3 Omni Flash Realtime 面向连续的音视频交互。它可以接收文字、图片、流式音频和视频画面,生成文字与语音回复,适合视觉问答、互动教学和口语陪练。例如用户展示一件物品并提问,模型结合当前画面回答;练习外语时,则可以围绕给定主题进行简短对话。

系统提示可以设定角色、使用语言与回答长度。为语音场景安排较短的句子,一次解释一个要点,再等待用户追问,通常更便于听取和继续交流。需要理解画面时,指出希望关注的物体或区域,并让模型对看不清的内容先追问。

实时交互怎样接入?

使用 WebSocket 建立会话后,先配置输出模态、音色、音频格式和轮次检测,再持续发送音频。视频通过连续画面参与理解,客户端应按协议提交图像,并保持音频与画面的对应关系。它使用实时事件流程,普通聊天请求不能直接替代这一过程。

为什么说完话还没有答复?

检查语音结束检测是否生效,以及输入是否真正提交。使用自动检测时,环境噪声或过短停顿可能影响轮次判断;手动模式则需要由客户端提交音频并触发回复。收到回复后还要正确解码、缓冲和播放音频,文字出现不代表扬声器已经有声音。

这一代实时模型侧重对话与多模态理解,不提供函数调用或内置联网搜索。需要查询业务系统时,由应用获取信息并传入会话。对画面外的内容不要做确定判断,可通过补充镜头或文字继续确认。

Use cases and prompting

先用纯语音验证一次完整对话,再加入画面。通过 session.update 配置文字和音频输出、支持的音色及音频格式,使用 input_audio_buffer.append 连续发送音频;手动结束时提交音频缓冲并触发 response.create。

可设置角色:“你是博物馆展品讲解员。结合用户当前画面,一次用两三句话解释一个特点;看不清展签时先请用户靠近,不猜测名称或年代。”

随后发送对应画面并口头提问。检查轮次是否正确结束、文字是否符合画面、音频能否正常播放。模型开始回答后,客户端应持续处理音频事件,避免只显示文字而丢弃语音数据。

API access

API documentation

qwen3-omni-flash-realtime

Use the linked documentation for model-specific request parameters and examples.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about qwen3-omni-flash-realtime

What is qwen3-omni-flash-realtime?

Compare qwen3-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call qwen3-omni-flash-realtime?

Create an API key with access to qwen3-omni-flash-realtime, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is qwen3-omni-flash-realtime priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of qwen3-omni-flash-realtime?

The model catalog lists a context window of 65536 tokens. Check the selected endpoint for request limits.

What is the maximum output of qwen3-omni-flash-realtime?

The model catalog lists a maximum output of 16384 tokens. Your request settings may set a lower limit.

How should I evaluate qwen3-omni-flash-realtime for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.