DeepSeek

DeepSeek V4 Flash 0731

DeepSeekToken-based
Alias:deepseek-v4-flash-0731
Create API Key

Compare deepseek-v4-flash-0731 API pricing, supported endpoints, capabilities and access options on Modelsell.

textfunction_callingtoolsjson_modestructured_outputreasoningcontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
943.7K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$0.44/ 1M
Output$1.33/ 1M
Cache Read$0.015/ 1M
Volcengine
-8%
火山引擎官方接口
Input$0.44$0.4048/ 1M
Output$1.33$1.2236/ 1M
Cache Read$0.015$0.0138/ 1M
火山引擎特价
-55%
火山引擎官方接口,高并发,活动特价
Input$0.44$0.198/ 1M
Output$1.33$0.5985/ 1M
Cache Read$0.015$0.00675/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsJSON modeStructured outputReasoning
Input
Output

Provider & data privacy

Provider
DeepSeekDocs
Tokenizer
DeepSeek tokenizer (BPE)
License
DeepSeek LicenseOpen weights
Data retention82 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·DeepSeek V4 Flash 0731

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About DeepSeek V4 Flash 0731

以编程任务为中心的 Flash 版本

DeepSeek V4 Flash 0731 面向代码、推理与智能体工作,适合在真实开发环境中阅读项目、修改实现和根据执行结果继续排错。这次正式版本强化了智能体能力,可以参与需要连续工具操作的工程任务,而不只是生成孤立的代码片段。

它处理文字输入并输出文字和代码。适合提供目录、相关源码、接口契约与测试结果,要求模型沿着具体功能追踪问题。例如,检查一次数据库迁移为什么导致旧客户端失效,或为已有项目补齐端到端功能并验证错误路径。图片中的内容应先转成文字资料,不能按视觉模型使用。

百万 Token 级上下文有利于阅读较大的资料集合。开发任务仍应说明入口、修改边界和验收方式,让它先建立依赖关系,再分阶段工作。对于长时间执行的任务,可以保留已完成步骤、失败原因和下一步验证目标,避免反复探索同一问题。

它提供 low、high 和 max 三档推理强度,适合按排错难度和任务预算选择。函数调用与结构化输出可以衔接搜索代码、读取文件、运行命令等工具;实际执行依赖开发环境提供这些能力。测试通过、程序行为符合预期和差异审查,才是判断开发结果的依据。

工程应用问题

适合处理跨文件修改吗? 可以给出一个完整功能目标,要求追踪数据与调用关系,再列出改动涉及的模块和回归范围。

什么时候用 max? 可用于较难的代码与工具任务;简单说明或局部修改可以先用较低强度,对比完成质量与耗时。

Use cases and prompting

给出真实项目里的完成条件

编程请求包含入口、限制与检查方法。

为这个命令行工具增加断点续传。
先阅读下载状态保存和重试逻辑,设计能兼容旧状态文件的最小改动。
实现后验证中断、恢复、文件损坏和重复启动四种情况,运行相关测试。
不要改变现有命令参数;最终给出修改文件、测试结果与尚存限制。

工具定义应明确参数与失败返回;将测试日志回传给模型继续分析。复杂任务可以提高 reasoning_effort,并为推理与最终答复留出预算;长任务在阶段结束时保存简短进度。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
officialUnlimitedUnlimitedUnlimited
VolcengineUnlimitedUnlimitedUnlimited
火山引擎特价UnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about deepseek-v4-flash-0731

What is deepseek-v4-flash-0731?

Compare deepseek-v4-flash-0731 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call deepseek-v4-flash-0731?

Create an API key with access to deepseek-v4-flash-0731, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is deepseek-v4-flash-0731 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of deepseek-v4-flash-0731?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of deepseek-v4-flash-0731?

The model catalog lists a maximum output of 943718 tokens. Your request settings may set a lower limit.

How should I evaluate deepseek-v4-flash-0731 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.