Gemini

Gemini 3.7 Flash

GoogleToken-based
Alias:gemini-3.7-flash
Create API Key

Compare gemini-3.7-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Mar 2026
Released
Aug 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.75$0.375/ 1M
Output$3.75$1.875/ 1M
Cache Read$0.075$0.0375/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.75$0.525/ 1M
Output$3.75$2.625/ 1M
Cache Read$0.075$0.0525/ 1M
Gemini Cli
-90%
Gemini cli 号池,适合 vibe coding学习等场景
Input$0.75$0.075/ 1M
Output$3.75$0.375/ 1M
Cache Read$0.075$0.0075/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention75 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.7 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.7 Flash

看懂过程,再推进工作

Gemini 3.7 Flash 是 Google 的多模态推理模型,能接收文字、图片、音频、视频和 PDF,并输出文本结果。它适合处理包含过程信息的材料,例如产品演示、培训录像、会议记录和复杂操作录屏,把“发生了什么”进一步整理为“接下来应该做什么”。

视频分析是它的重要应用方向。可以要求模型沿着某个目标追踪事件:演示中哪一步引入了错误,用户在哪些环节停顿,讲解内容与画面是否一致。把视频和操作手册一并提供,还可以让它逐条比对实际流程与书面要求,列出差异及对应片段。结果需要便于复查时,应要求记录时间位置和画面依据。

它也能处理编程与企业流程任务。结合代码、需求文档和工具返回的信息,可以拆解实现步骤、定位问题或整理交付物。官方能力包括搜索辅助、代码执行、函数调用和结构化输出;在应用开放这些工具后,模型可以依据搜索结果、计算结果或业务接口反馈继续推理。

思考强度可按任务选择低、中、高。简单整理可以减少思考投入,跨视频与文档的复杂对照则适合增加推理空间。它输出的是文本分析,不直接生成音频或图片。对于细小文字、快速变化的画面和含糊的音频,补充清晰片段与背景说明,会更有利于得到可靠结论。

Use cases and prompting

视频与手册对照

上传演示视频和操作手册,说明要检查的流程。

提示词示例: “请检查这段设备演示是否符合附件中的操作顺序。按手册步骤列出视频中的对应时间位置、实际动作和差异。画面无法确认的步骤标为待复查,最后生成一份供培训人员使用的修改清单。”

常见问题

  • 视频很长,怎样避免只得到概述?指定事件清单和时间粒度,先建立片段索引,再深入分析有疑问的部分。
  • 能一边分析一边算数据吗?应用启用代码执行后,可以让模型用计算结果支持判断,并保留输入数据与计算说明。
  • minimal 思考设置报错怎么办?该模型支持低、中、高思考强度,选择其中一种。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.7-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Gemini CliUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.7-flash

What is gemini-3.7-flash?

Compare gemini-3.7-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.7-flash?

Create an API key with access to gemini-3.7-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.7-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.7-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.7-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.7-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.