DeepSeek

DeepSeek V4 Flash

DeepSeekToken-based
Alias:deepseek-v4-flash
Create API Key

Compare deepseek-v4-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textfunction_callingtoolsjson_modestructured_outputreasoningcontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
943.7K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.44/ 1M
Output$1.33/ 1M
Cache Read$0.015/ 1M
Volcengine
-8%
火山引擎官方接口
Input$0.44$0.4048/ 1M
Output$1.33$1.2236/ 1M
Cache Read$0.015$0.0138/ 1M
Alibaba
阿里巴巴百炼官方
Input$0.44/ 1M
Output$1.33/ 1M
Cache Read$0.015/ 1M
火山引擎特价
-55%
火山引擎官方接口,高并发,活动特价
Input$0.44$0.198/ 1M
Output$1.33$0.5985/ 1M
Cache Read$0.015$0.00675/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoning
Input
Output

Provider & data privacy

Provider
DeepSeekDocs
Tokenizer
DeepSeek tokenizer (BPE)
License
DeepSeek LicenseOpen weights
Data retention88 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·DeepSeek V4 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About DeepSeek V4 Flash

适合频繁交互的文本助手

DeepSeek V4 Flash 面向注重响应效率的推理与编程任务。它适合开发者在写代码时快速讨论实现,客服或内部知识助手根据资料回答问题,也适合在智能体流程中反复整理信息、选择下一步并处理工具返回结果。

速度之外,保留问题的上下文

模型可以处理较长的文本材料。使用知识助手时,可以提供相关制度、产品说明或项目文档,让回答围绕真实材料展开;使用编程助手时,可以给出当前函数、调用方式和错误信息,让修改建议落到具体实现上。先限定问题范围,通常比一次提出许多不相关要求更有效。

Flash适合快速形成候选方案,再通过反馈继续调整。可以让它先指出最可能的原因、提出一个最小改动,然后根据实际结果决定是否继续。若问题涉及多层约束或材料冲突,应保留必要的分析步骤,不能只追求尽快输出答案。

把工具结果接回对话

应用提供工具调用后,模型可以参与查询、代码执行和自动化处理。工具失败或返回空结果时,应如实保留,避免模型用推测填补。固定任务样本评估准确率、完整性和响应表现,有助于判断它是否适合你的交互场景。需要引用资料的问答,最好让回答带上对应段落或文件位置,方便用户继续查看。

Use cases and prompting

内部知识助手示例

“仅根据下方产品手册回答:管理员怎样撤销成员权限?请给出操作步骤和相关段落位置。如果手册没有说明撤销后的数据归属,明确标为待确认,不要自行补全。答案控制在可直接发给同事的长度。”

多轮追问时保留当前产品版本与已经引用的材料。

编程问题怎样问得更有效?

给出相关函数、输入示例、实际错误和预期结果,先要求最小修复,再根据测试结果继续。

知识助手为什么回答了资料外的内容?

明确限定依据,设置材料不足时的返回方式,并要求保留出处。工具没有查到信息时,应用也应把空结果交回模型。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
VolcengineUnlimitedUnlimitedUnlimited
火山引擎特价UnlimitedUnlimitedUnlimited
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited
user_owned:1366:188UnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about deepseek-v4-flash

What is deepseek-v4-flash?

Compare deepseek-v4-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call deepseek-v4-flash?

Create an API key with access to deepseek-v4-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is deepseek-v4-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of deepseek-v4-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of deepseek-v4-flash?

The model catalog lists a maximum output of 943718 tokens. Your request settings may set a lower limit.

How should I evaluate deepseek-v4-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.