XiaomiMiMo

MiMo V2.6 Flash

XiaomiToken-based
Alias:xiaomi/mimo-v2.6-flash
Create API Key

Compare xiaomi/mimo-v2.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagevideoaudiofunction_callingtoolsjson_modestructured_outputreasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1.1M
Maximum input window
Max output
131.1K
Maximum tokens per response
Modalities
→

Pricing by Supplier

XiaomiMIMO
XiaomiMIMO 官方
Input$75/ 1M
Output$75/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
Unknown
Tokenizer
BPE (vendor-specific)
License
Provider-specificUnknown
Data retention25 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·MiMo V2.6 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About MiMo V2.6 Flash

把录音、画面与文字一起读懂

MiMo-V2.6 Flash 适合日常专业工作中的多模态内容处理。它可以联合理解文字、图片、视频和音频,再输出文字结果。例如结合讲解录音与演示画面提炼要点,阅读商品图片和说明整理属性,或把操作录像中的过程写成步骤清单。用户不必把每一种材料都手工转述成文字后再分析。

处理会议或培训资料时,可以给模型明确目标:区分已确认决定与讨论建议,找出负责人、待办和未解决问题,再将结果按主题组织。多份材料存在差异时,要求指出证据所在,避免把画面中出现的文字直接当成发言者已经确认的结论。

怎样用于代码和连续工作流?

模型支持深度思考与工具调用,可根据需求分析代码、提出实现方案,并借助应用提供的工具读取文件、执行检查或查询数据。适合把工作拆成小步:先定位相关材料,再处理一项修改,获得结果后继续。工具是否成功应以执行返回为准,并将新结果纳入后续对话。

百万 Token 级上下文可容纳较长材料,但建议按文件、时间段与主题分组,明确当前要解决的问题。简单整理可使用直接回复方式,涉及跨材料推断或多步代码分析时再启用思考,并为正式结论留出输出空间。

输出能接入业务系统吗?

模型可生成结构化结果,适合内容标签、会议行动项和资料归档。提示中写清字段、缺失值和重复项处理方式,应用再校验并保存。输出为文字;若要播放语音,需要另外接入语音合成服务。

Use cases and prompting

例如整理产品培训,提交录音、演示视频和补充文字后输入:

“请形成一份面向新同事的培训笔记。按‘功能用途、演示步骤、注意事项、尚未回答的问题’整理。区分讲解者明确说明的规则与画面中的临时示例;每项结论标注材料名称和可定位的片段,无法确定时写待确认。最后提取行动项,包含任务、负责人、截止日期,未提及的字段留空。”

长录音与视频可先按主题切段,再汇总笔记;保留片段编号与原始时间范围,便于回看。需要写入任务系统时,先校验结构化字段,并确认负责人和日期来自材料。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
XiaomiMIMOUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about xiaomi/mimo-v2.6-flash

What is xiaomi/mimo-v2.6-flash?

Compare xiaomi/mimo-v2.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call xiaomi/mimo-v2.6-flash?

Create an API key with access to xiaomi/mimo-v2.6-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is xiaomi/mimo-v2.6-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of xiaomi/mimo-v2.6-flash?

The model catalog lists a context window of 1050000 tokens. Check the selected endpoint for request limits.

What is the maximum output of xiaomi/mimo-v2.6-flash?

The model catalog lists a maximum output of 131072 tokens. Your request settings may set a lower limit.

How should I evaluate xiaomi/mimo-v2.6-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.