Gemini

Gemini 3.6 Flash

GoogleToken-based
Alias:gemini-3.6-flash
Create API Key

Compare gemini-3.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Mar 2026
Released
Jul 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$1.5$0.75/ 1M
Output$7.5$3.75/ 1M
Cache Read$0.15$0.075/ 1M
Google Vertex
-30%
谷歌官方接口
Input$1.5$1.05/ 1M
Output$7.5$5.25/ 1M
Cache Read$0.15$0.105/ 1M
Gemini Cli
-90%
Gemini cli 号池,适合 vibe coding学习等场景
Input$1.5$0.15/ 1M
Output$7.5$0.75/ 1M
Cache Read$0.15$0.015/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention30 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.6 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.6 Flash

让分析跟上迭代节奏

Gemini 3.6 Flash 面向真实工作中的快速分析与执行循环,擅长代码生成、智能体协作和空间关系理解。它适合需要连续读取信息、提出改动、查看结果再修正的项目,也适合将多种材料放在同一任务里综合判断。

不同素材可以围绕同一个问题

模型支持文字、图片、视频、音频和PDF输入,输出文本。比如分析一次产品使用测试,可以结合操作录像、用户讲述与界面说明,整理卡点发生的位置和原因;审查实现时,可以对照截图、设计要求和代码,指出需要调整的逻辑或布局。素材的价值在于互相补充,而不是只把它们各自摘要一遍。

空间关系任务应明确观察目标,例如物体位于哪里、路径如何变化、画面中哪个元素遮挡了另一个元素。涉及精细位置和快速动作时,应让结论对应具体画面,并把看不清的部分保留为待确认。

配合工具逐步完成工作

通过已接入的函数调用、代码执行或检索工具,模型可以参与持续的开发与资料处理流程。把每轮目标控制在可检查的范围内,及时提供测试输出、错误信息或新截图,能让后续修改更有依据。需要近期事实时应结合实时资料;处理影音输入时,结果是文字分析,配音和视频生成需要另行安排。

Use cases and prompting

使用测试分析示例

“结合操作录像、用户访谈音频和页面说明,找出用户从选商品到付款时的阻碍。按发生顺序列出问题,注明对应画面或话语,区分观察到的事实与推测原因。最后提出三项最小改动,并说明怎样验证改善效果。”

素材较长时,先指定需要关注的流程或时间段。

如何用于快速编程循环?

提供当前代码、一个明确目标和可执行的验证方式。每轮把实际测试结果交回,让模型根据失败原因修正,而不是重新生成整套方案。

上传音频或视频后会生成新的媒体吗?

该模型输出文字,适合理解和分析影音内容;制作新视频或语音应使用相应生成模型。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.6-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
Gemini CliUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.6-flash

What is gemini-3.6-flash?

Compare gemini-3.6-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.6-flash?

Create an API key with access to gemini-3.6-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.6-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.6-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.6-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.6-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.