XiaomiMiMo

MiMo V2.5 Tts Voicedesign

XiaomiToken-based
Alias:mimo-v2.5-tts-voicedesign
Create API Key

Compare mimo-v2.5-tts-voicedesign API pricing, supported endpoints, capabilities and access options on Modelsell.

tts语音合成风格控制声音设计
Starting price
Input / Output · 1M
Context
8.2K
Maximum input window
Max output
8.2K
Maximum tokens per response
Modalities
→

Pricing by Supplier

XiaomiMIMO
XiaomiMIMO 官方
Free

Capabilities / Supported modalities

Function callingToolsJSON modeStructured output
Input
Output

Provider & data privacy

Provider
OpenAIDocs
Tokenizer
cl100k_baseOlder GPT-3.5 family
License
Proprietary (commercial)Proprietary
Data retention30 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·MiMo V2.5 Tts Voicedesign

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About MiMo V2.5 Tts Voicedesign

用几句话描述你想要的声音

MiMo-V2.5-TTS Voicedesign 适合没有录音样本、希望从文字出发创造声音的任务。你可以描述年龄感、声线质地、情绪和节奏,生成对应方向的语音,用于原创角色、纪录片旁白、播客开场或品牌声音的早期探索。

声音描述放在 user 消息,这是本型号的必要输入;待朗读正文通常放在 assistant 消息。描述不必很长,一到四句即可,重点是核心特征。例如“中年男声,略沙哑但清晰,语速缓慢,像在讲述长途旅行的见闻”,比“普通、好听的声音”更容易形成明确方向。

可以加上角色身份、说话场景和表达方式,让声音与内容相互配合。平静的睡前故事适合柔和节奏,紧张的比赛解说适合更强的情绪推进。避免同时要求互相冲突的特征,也不要把混响、均衡器或压缩等后期效果当作音色描述。

这个型号根据文字设计声音,不使用内置 voice,也不依据录音复刻,且不支持唱歌模式。可选 optimize_text_preview 用于智能润色播报正文;如果必须逐字朗读,应先明确文本处理要求并核对结果。低延迟流式输出尚未开放,兼容流式接口会在全部推理完成后一次返回结果。

设计声音的问题

怎样比较不同设计方向? 保留同一段试听正文,每次调整少量特征,比较清晰度、情绪和角色适配程度。

能保证长系列每段声音完全一样吗? 需要实际试听验证一致性,持续使用相同描述,并检查分段衔接和表达变化。

Use cases and prompting

将音色描述与正文分开

user:成年女性旁白,声线清晰略带颗粒感,温暖而从容。语速适中,像在安静地介绍自然纪录片,不使用夸张播音腔。
assistant:清晨的第一束光穿过树叶,森林里的声音渐渐醒来。

先生成短段试听,再调整音色、情绪或节奏。不要传参考录音或内置音色名称。需要原文不变时谨慎使用 optimize_text_preview,并回听核对;按整段结果等待、保存与播放,长内容分段检查声音和音量衔接。

API access

API documentation

mimo-v2.5-tts-voicedesign

Use the linked documentation for model-specific request parameters and examples.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Rate limits

SupplierRPMTPMRPD
XiaomiMIMOUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about mimo-v2.5-tts-voicedesign

What is mimo-v2.5-tts-voicedesign?

Compare mimo-v2.5-tts-voicedesign API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call mimo-v2.5-tts-voicedesign?

Create an API key with access to mimo-v2.5-tts-voicedesign, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is mimo-v2.5-tts-voicedesign priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of mimo-v2.5-tts-voicedesign?

The model catalog lists a context window of 8192 tokens. Check the selected endpoint for request limits.

What is the maximum output of mimo-v2.5-tts-voicedesign?

The model catalog lists a maximum output of 8192 tokens. Your request settings may set a lower limit.

How should I evaluate mimo-v2.5-tts-voicedesign for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.