Gemini

Gemini 3.5 Flash

GoogleToken-based
Alias:gemini-3.5-flash
Create API Key

Compare gemini-3.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Jan 2025
Released
May 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$1.5$0.75/ 1M
Output$9$4.5/ 1M
Cache Read$0.15$0.075/ 1M
Google Vertex
-30%
谷歌官方接口
Input$1.5$1.05/ 1M
Output$9$6.3/ 1M
Cache Read$0.15$0.105/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention5 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.5 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.5 Flash

将复杂工作拆成能推进的步骤

Gemini 3.5 Flash 适合智能体协作、多步骤流程和需要持续迭代的编程任务。它可以围绕明确目标读取材料、整理信息、提出行动,再结合实际结果继续修正。对于需要处理大量资料或并行子任务的应用,它也适合承担分工明确的一部分。

多模态输入帮助补齐信息

模型接受文本、图片、视频、音频与PDF,输出文本。比如制作培训资料时,可以把讲课录音、演示录像和讲义结合起来,整理讲解顺序、关键操作和学员容易遗漏的步骤;处理产品项目时,可以让它对照需求、界面截图和反馈记录,形成可执行的修改清单。

把长流程拆成资料提取、问题判断和成果整理等阶段,每阶段给出清楚输入与返回格式,能更容易追踪工作是否完成。子任务应保留出处和必要上下文,后续汇总才有依据,而不是只收到彼此孤立的结论。

工具协作中的快速反馈

通过应用接入函数调用、代码执行和检索等工具后,模型可以参与开发与资料处理循环。测试失败、检索缺失和材料冲突都应回传,作为下一轮工作的输入。长任务要保存中间成果与未完成事项,短任务则控制在一个可检查目标内,避免为了一个简单修改不断扩展范围。

Use cases and prompting

培训资料整理示例

“根据课程录音、操作录像和讲义,整理一份新员工培训稿。先提取实际演示过的步骤,再标出讲义中未演示的内容。按准备、操作、检查三个阶段组织,每一步说明输入、动作和完成标志。材料有冲突时保留原说法并列出待确认项。”

怎样用于多个子任务?

为每个子任务给出独立目标和一致的返回格式,例如一项提取操作步骤,另一项整理问题与答案;汇总时保留来源信息。

长流程做到一半偏题怎么办?

在每阶段结束时对照原目标检查成果,保存已完成步骤与下一步。新要求要明确优先级,避免把背景材料里的每个话题都扩成新任务。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.5-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.5-flash

What is gemini-3.5-flash?

Compare gemini-3.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.5-flash?

Create an API key with access to gemini-3.5-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.5-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.5-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.5-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.5-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.