MoonshotAI

Kimi K3

MoonshotToken-based
Alias:kimi-k3
Create API Key

Compare kimi-k3 API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagevideofunction_callingtoolsjson_modestructured_outputreasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
943.7K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$2.86/ 1M
Output$14.29/ 1M
Cache Read$0.286/ 1M
Moonshot
Kimi 官方渠道
Input$2.86/ 1M
Output$14.29/ 1M
Cache Read$0.286/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
Moonshot AIDocs
Tokenizer
Kimi tokenizer
License
Proprietary (commercial)Proprietary
Data retention7 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Kimi K3

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Kimi K3

从资料理解走向成果制作

Kimi K3 是面向长周期编程、知识工作和复杂推理的多模态模型,可以理解文字、图片和视频。它适合需要持续处理大量上下文的项目:读懂大型代码库,围绕多份材料开展研究,或将分析结果发展成报告、可视化方案与交互界面。

让视觉反馈参与修改

在开发工作中,可以让模型结合需求、代码和运行截图判断实现是否到位;在研究工作中,可以提供图表、视频资料与文字记录,让它寻找共同证据和冲突信息。视觉输入的价值是帮助模型观察实际表现,后续改动仍应通过新的截图或工具结果检查。

K3适合长时间保持一个明确目标。项目可以分为理解材料、建立方案、制作成果和验证修正等阶段,每阶段保存关键决策与待解决问题。面对大型仓库,先建立文件与模块索引,再分析具体路径;面对研究材料,保留来源与判断依据,避免结论脱离原文。

工具决定实际交付形式

模型输出文本;应用提供终端、文件和浏览器等工具后,可以进一步参与代码实现、图表制作和交互成果的生成。要求它交付可检查的文件或界面,并实际打开或运行验证,才能判断工作完成度。复杂任务中的自主推进也需要明确范围,尤其说明哪些已有行为、数据与视觉约定必须保留。

Use cases and prompting

研究看板示例

“根据用户访谈、功能使用数据和产品演示,分析新用户流失发生在哪一步。先整理证据与假设,再设计一个可交互的分析看板:展示各阶段人数、关键问题和验证计划。所有图表注明计算方式,缺失数据保留为空,不要补造数值。”

需要真实看板时,提供数据处理、文件和预览工具。

怎样把长研究任务接到可用成果?

先明确读者和决策目标,再规定交付形式及检查方式;让图表、结论与原始数据对应。

视觉反馈应该在什么时候加入?

在初版完成后提供实际截图或录像,让模型检查布局和操作效果,再依据观察继续修正。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
MoonshotUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about kimi-k3

What is kimi-k3?

Compare kimi-k3 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call kimi-k3?

Create an API key with access to kimi-k3, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is kimi-k3 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of kimi-k3?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of kimi-k3?

The model catalog lists a maximum output of 943718 tokens. Your request settings may set a lower limit.

How should I evaluate kimi-k3 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.