DeepSeek

DeepSeek V4.1 Flash

DeepSeekToken-based
Alias:deepseek-v4.1-flash
Create API Key

Compare deepseek-v4.1-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
393.2K
Maximum tokens per response
Modalities
→

Pricing by Supplier

official
官方接口直连
Input$75/ 1M
Output$75/ 1M

Capabilities / Supported modalities

ReasoningFunction callingToolsJSON modeVision
Input
Output

Provider & data privacy

Provider
DeepSeekDocs
Tokenizer
DeepSeek tokenizer (BPE)
License
DeepSeek LicenseOpen weights
Data retention83 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·DeepSeek V4.1 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About DeepSeek V4.1 Flash

从文本推理走向视觉协作

DeepSeek V4.1 Flash 可以结合文字与图片进行分析,适合编程助手、终端工作流和计算机使用智能体。它能围绕代码、日志与界面截图理解一个问题,让视觉信息参与后续判断,而不是把截图内容与文本任务分开处理。

适合需要多次观察和行动的工作

例如排查一个后台页面,可以先读取错误日志,再对照截图判断用户实际看到的状态,随后修改代码并检查结果。长流程任务则可以保留项目材料、关键约束与阶段成果,让模型持续推进到有明确完成标志的结果。应用提供终端、浏览器或其他工具后,这些分析才能转成实际操作。

图像理解适合界面分析、图表解读和视觉信息提取。输入截图时,应保留必要的上下文,并说明希望发现的问题;文字很小或画面不完整时,补充局部清晰截图比要求模型猜测更有效。

长任务也需要明确边界

V4.1 Flash支持思考与非思考模式,也支持JSON输出和工具调用。可以按任务选择相应方式:格式化提取重在字段稳定,复杂排障重在证据链与验证。多步骤执行中要保留失败信息和中间结果,并设置阶段检查点,避免模型在缺少实际状态时继续推测。

Use cases and prompting

视觉排障示例

“这是后台保存设置后的截图、浏览器报错和对应组件代码。请先说明界面状态与日志能共同证明什么,再沿提交与回显流程定位问题。提出最小修复,并验证保存失败、重复提交和刷新后回显三种情况。”

需要实际修改与测试时,让应用提供相关文件和执行工具。

截图里能看见信息,为什么仍会识别错?

小字、压缩和遮挡会影响判断。提供原尺寸或局部放大的清晰截图,并说明要关注的区域。

怎样让多步骤任务有明确结束?

事先写出完成条件,例如测试通过、页面回显正确、输出文件可打开;每一步根据实际结果检查,而不是只确认计划已写完。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about deepseek-v4.1-flash

What is deepseek-v4.1-flash?

Compare deepseek-v4.1-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call deepseek-v4.1-flash?

Create an API key with access to deepseek-v4.1-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is deepseek-v4.1-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of deepseek-v4.1-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of deepseek-v4.1-flash?

The model catalog lists a maximum output of 393216 tokens. Your request settings may set a lower limit.

How should I evaluate deepseek-v4.1-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.