Gemini

Gemini 3 Flash

GoogleToken-based
Alias:gemini-3-flash
Create API Key

Compare gemini-3-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Released
Dec 2025

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.5$0.25/ 1M
Output$3$1.5/ 1M
Cache Read$0.05$0.025/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.5$0.35/ 1M
Output$3$2.1/ 1M
Cache Read$0.05$0.035/ 1M
Gemini Cli
-90%
Gemini cli 号池,适合 vibe coding学习等场景
Input$0.5$0.05/ 1M
Output$3$0.3/ 1M
Cache Read$0.05$0.005/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention73 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3 Flash

从图文、录音和视频中理解任务

Gemini 3 Flash 适合内容来源多样、又需要及时响应的助手应用。它可以阅读文档与图片、理解音频和视频,再用文字回答问题、整理信息或生成代码。教育内容分析、产品演示解读、图文客服和交互原型开发,都可以把它作为工作流程的一部分。

例如,给它一段软件操作视频,要求总结用户完成任务的路径,并标出界面卡住的位置;给它一份 PDF 和配套图表,要求解释结论是否得到数据支持;给它培训录音,要求生成复习提纲和练习题。这些任务都需要在请求中说明关注点,而不仅是要求“总结一下”。

在开发场景中,它能根据产品说明和视觉参考生成界面方案、交互逻辑与代码。结合真实的文件编辑、运行和电脑操作工具,还可以参与检查页面效果的多步骤工作。工具需要在应用中配置,输出代码后也应实际运行并检查交互。

它支持思考、函数调用、结构化输出,以及按需启用的搜索 grounding、URL 上下文和代码执行。较大的上下文便于关联多份材料,返回 JSON 则适合让系统读取提取结果。媒体分析结果以文字输出,生成图片、语音或视频需要选择相应生成服务。

如何提问

能从视频里找到某个操作吗? 可以指定要查找的动作,要求返回相关时间位置及观察依据,便于重新检查原片。

适合做交互原型吗? 可以描述页面目标、用户路径和数据状态,要求生成原型代码与需要验证的交互清单,再在开发环境中运行。

Use cases and prompting

围绕观察结果给出任务

提交媒体时说明要回答的具体问题,并约定定位方式。

请分析这段购物流程录像。
找出用户从商品页到支付页遇到的重复输入、等待和错误提示。
每项记录时间点、界面状态、影响与改进建议。
最后按优先级给出三项修改,并生成相关交互的验收步骤;不要推测视频未展示的后台行为。

生成界面时提供技术栈、参考截图和需要保留的组件。结构化提取时定义字段;需要最新信息时启用搜索,并要求将检索到的事实与自己的分析分开。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Gemini CliUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3-flash

What is gemini-3-flash?

Compare gemini-3-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3-flash?

Create an API key with access to gemini-3-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.