qwen3-omni-flash-realtimeCompare qwen3-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.
Qwen tokenizer (tiktoken-compat)Scores on standardized evaluations. Higher percentages are better — and rank percentile shows
Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3 Omni Flash Realtime
This model is not included in the current benchmark snapshot.
Missing models or measurements are not zero scores.
Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.
Qwen3 Omni Flash Realtime 面向连续的音视频交互。它可以接收文字、图片、流式音频和视频画面,生成文字与语音回复,适合视觉问答、互动教学和口语陪练。例如用户展示一件物品并提问,模型结合当前画面回答;练习外语时,则可以围绕给定主题进行简短对话。
系统提示可以设定角色、使用语言与回答长度。为语音场景安排较短的句子,一次解释一个要点,再等待用户追问,通常更便于听取和继续交流。需要理解画面时,指出希望关注的物体或区域,并让模型对看不清的内容先追问。
使用 WebSocket 建立会话后,先配置输出模态、音色、音频格式和轮次检测,再持续发送音频。视频通过连续画面参与理解,客户端应按协议提交图像,并保持音频与画面的对应关系。它使用实时事件流程,普通聊天请求不能直接替代这一过程。
检查语音结束检测是否生效,以及输入是否真正提交。使用自动检测时,环境噪声或过短停顿可能影响轮次判断;手动模式则需要由客户端提交音频并触发回复。收到回复后还要正确解码、缓冲和播放音频,文字出现不代表扬声器已经有声音。
这一代实时模型侧重对话与多模态理解,不提供函数调用或内置联网搜索。需要查询业务系统时,由应用获取信息并传入会话。对画面外的内容不要做确定判断,可通过补充镜头或文字继续确认。
先用纯语音验证一次完整对话,再加入画面。通过 session.update 配置文字和音频输出、支持的音色及音频格式,使用 input_audio_buffer.append 连续发送音频;手动结束时提交音频缓冲并触发 response.create。
可设置角色:“你是博物馆展品讲解员。结合用户当前画面,一次用两三句话解释一个特点;看不清展签时先请用户靠近,不猜测名称或年代。”
随后发送对应画面并口头提问。检查轮次是否正确结束、文字是否符合画面、音频能否正常播放。模型开始回答后,客户端应持续处理音频事件,避免只显示文字而丢弃语音数据。
Use the linked documentation for model-specific request parameters and examples.
All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.
Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.
| Supplier | RPM | TPM | RPD |
|---|---|---|---|
| Alibaba | Unlimited | Unlimited | Unlimited |
| official | Unlimited | Unlimited | Unlimited |
No restriction
Compare qwen3-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.
Create an API key with access to qwen3-omni-flash-realtime, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.
Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.
The model catalog lists a context window of 65536 tokens. Check the selected endpoint for request limits.
The model catalog lists a maximum output of 16384 tokens. Your request settings may set a lower limit.
Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.
