Gemini

Gemini 3.8 Flash

GoogleToken-based
Alias:gemini-3.8-flash
Create API Key

Compare gemini-3.8-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Released
Sep 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.75$0.375/ 1M
Output$3.75$1.875/ 1M
Cache Read$0.075$0.0375/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.75$0.525/ 1M
Output$3.75$2.625/ 1M
Cache Read$0.075$0.0525/ 1M
Gemini Cli
-90%
Gemini cli 号池,适合 vibe coding学习等场景
Input$0.75$0.075/ 1M
Output$3.75$0.375/ 1M
Cache Read$0.075$0.0075/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention55 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.8 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.8 Flash

面向持续执行的复杂流程

Gemini 3.8 Flash 适合长周期软件工程、自主智能体和复杂企业工作流。它可以围绕一个目标连续处理资料、规划步骤、分析结果并修正方案,适合将代码开发、业务规则与实际操作连接起来的项目。

让不同材料进入同一工作上下文

模型支持文字、图片、视频、音频和PDF输入,输出文本。代码任务可以结合仓库说明、错误日志和界面录像;企业流程任务可以结合操作手册、系统截图和会议记录。把这些信息放在同一目标下,有助于识别实现与流程之间的缺口,而不是孤立地解释每份材料。

长任务应明确必须满足的约束和阶段成果,例如接口保持兼容、数据处理可回退、异常状态有记录。模型接入工具后可以参与代码执行、函数调用和检索,应用应把每次真实结果交回,让后续工作依据实际状态推进。

为长流程设置检查点

复杂任务适合分为理解现状、设计改动、实施验证和整理交付等阶段。每一阶段都保存已完成内容、失败原因和下一步,减少重复分析,也方便加入新的限制。Gemini 3.8 Flash 支持low、medium和high思考等级,可按任务选择。涉及不确定信息时应保留待确认项,不能把模型的自信程度当作验证结果。

Use cases and prompting

系统整合示例

“根据现有工单系统代码、操作手册和客服演示录像,设计自动分派流程。先整理当前状态和分派规则,再找出冲突与缺失信息。分步骤实现,保留人工改派入口,并验证重复请求、无匹配人员和处理中人员离线等情况。每阶段输出实际结果与剩余问题。”

如何防止长任务重复做已完成的工作?

保存阶段检查点,包含完成内容、关键决策、验证结果和待办。继续时先读取检查点,再处理下一项。

minimal思考设置为什么报错?

Gemini 3.8 Flash支持low、medium、high,不支持minimal。应选择受支持的等级,并根据实际质量与响应时间评估。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.8-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Gemini CliUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.8-flash

What is gemini-3.8-flash?

Compare gemini-3.8-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.8-flash?

Create an API key with access to gemini-3.8-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.8-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.8-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.8-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.8-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.