BaiLian

Qwen3.5 Omni Flash Realtime

AlibabaToken-based
Alias:qwen3.5-omni-flash-realtime
Create API Key

Compare qwen3.5-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.

audiotextimagevideorealtimestreaming多模态对话
Starting price
Input / Output · 1M
Context
262.1K
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Released
Mar 2026

Pricing by Supplier

official
官方接口直连
Input$75/ 1M
Output$75/ 1M
Alibaba
阿里巴巴百炼官方
Input$75/ 1M
Output$75/ 1M

Capabilities / Supported modalities

StreamingVisionSystem promptFunction callingToolsWeb search
Input
Output

Provider & data privacy

Provider
Alibaba (Qwen)Docs
Tokenizer
Qwen tokenizer (tiktoken-compat)
License
Tongyi Qianwen LicenseOpen weights
Data retention63 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Qwen3.5 Omni Flash Realtime

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Qwen3.5 Omni Flash Realtime

边看边聊的语音助手

Qwen3.5 Omni Flash Realtime 适合实时讲解、语音客服和互动教学。它能结合文字、图片、麦克风音频与视频画面理解问题,再以流式文字和语音回复。例如用户拍摄设备面板并口头描述操作,助手可以解释可见按钮与提示信息,再追问缺少的条件;也可以在展厅中根据镜头内容逐步介绍展品。

实时会话通过 WebSocket 传输输入与事件。应用持续发送音频片段和需要关注的画面,并接收增量语音播放。角色提示应明确回答语言、每轮长度和追问方式,让口头回复简短、容易跟随,长步骤分轮解释。

怎样让对话轮次更自然?

模型支持语义打断,semantic_vad 可结合语义有效性判断用户话语,过滤部分附和声与背景噪声。应用还应处理播放队列:用户转入新问题时及时停止旧回复,避免已经排队的语音继续播放。安静房间、商场噪声和用户中途停顿的表现都值得分别检查。

能查询业务信息或最新资料吗?

需要订单、预约等业务数据时,可配置函数定义,由模型提出调用,应用执行并传回真实结果,再生成口头答复。需要查询公开的最新资料时,可启用联网搜索。tools 与 enable_search 不能同时启用,应按会话用途选择;业务操作和搜索结果都应明确传回模型,避免把准备执行的步骤说成已经完成。

Use cases and prompting

建立 WebSocket 会话后,用 session.update 配置文字与音频输出、音色,以及 turn_detection.type="semantic_vad"。例如设置角色提示:

“你是设备使用助手。结合当前画面与用户语音,每轮最多解释两个步骤,再等用户确认。只描述能看清的标签;看不清时请用户靠近镜头。用户更换问题时立即跟随新问题。”

通过 input_audio_buffer.append 发送音频,按需要补充画面,接收 response.audio.delta 播放语音。接入自定义函数时,读取调用参数,执行后用 conversation.item.create 回传带相同 call_id 的 function_call_output,再触发 response.create。只使用联网搜索的会话可启用 enable_search=true,不要同时配置 tools。

API access

API documentation

qwen3.5-omni-flash-realtime

Use the linked documentation for model-specific request parameters and examples.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about qwen3.5-omni-flash-realtime

What is qwen3.5-omni-flash-realtime?

Compare qwen3.5-omni-flash-realtime API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call qwen3.5-omni-flash-realtime?

Create an API key with access to qwen3.5-omni-flash-realtime, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is qwen3.5-omni-flash-realtime priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of qwen3.5-omni-flash-realtime?

The model catalog lists a context window of 262144 tokens. Check the selected endpoint for request limits.

What is the maximum output of qwen3.5-omni-flash-realtime?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate qwen3.5-omni-flash-realtime for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.