Gemini

Gemini 2.5 Flash

GoogleToken-based
Alias:gemini-2.5-flash
Create API Key

Compare gemini-2.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideocachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Jan 2025
Released
Jun 2025

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.3$0.15/ 1M
Output$2.5$1.25/ 1M
Cache Read$0.03$0.015/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.3$0.21/ 1M
Output$2.5$1.75/ 1M
Cache Read$0.03$0.021/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention17 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 2.5 Flash

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 2.5 Flash

为高频应用保留分析能力

Gemini 2.5 Flash 适合需要及时响应、又不止做简单匹配的任务。它可以理解文字、图片、音频和视频,再输出文字答案、摘要、代码或结构化结果。客服知识问答、内容分析、会议资料整理和业务助手,都可以围绕它建立处理流程。

它支持思考,能够在快速处理与更深入分析之间按任务需要安排预算。例如,对普通消息提取诉求,对包含多个条件的投诉分析处理路径;对短视频生成摘要,对重点片段进一步解释动作和上下文。任务越明确,越容易兼顾响应速度与有用程度。

多模态输入适合把“看到了什么”和“应如何处理”联系起来。你可以给它设备照片与故障描述,让它整理可见现象和需要确认的检查项;也可以提供培训录像,让它按章节形成笔记、练习题和常见误解。涉及细节时,应要求标记图片区域或视频时间,便于回到原始材料复核。

函数调用与结构化输出让它可以衔接查询、分类和后续业务动作。需要外部资料或计算时,可通过调用端配置搜索、URL 上下文和代码执行。长上下文适合多份关联资料,但建议把关注问题放在前面,并为材料编号,避免摘要偏离重点。

常见场景

适合做客服助手吗? 可以结合知识库内容解释流程、识别诉求与生成回复;订单查询和实际退款应由业务工具完成。

能输出带声音的视频吗? 本模型输出文字,音视频输入用于理解;需要生成媒体时,应使用专门的生成模型。

Use cases and prompting

让多模态内容服务于具体问题

提交图片、录音或视频后,说明关注对象与期望输出。

这是设备操作教学视频和学员的三个问题。
为每个问题找出视频中的相关步骤,给出时间位置、操作说明和容易出错的地方。
再生成一份 6 步检查清单,用简短句子写给第一次操作的人。
没有在视频出现的安全要求不要自行补充为厂家规定。

客服或业务提取使用固定 JSON 字段,并区分原文事实和建议。需要更深入分析的请求可增加思考预算;简单摘要则约定长度与必含主题,减少无关内容。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-2.5-flash:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
Google VertexUnlimitedUnlimitedUnlimited
defaultUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-2.5-flash

What is gemini-2.5-flash?

Compare gemini-2.5-flash API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-2.5-flash?

Create an API key with access to gemini-2.5-flash, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-2.5-flash priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-2.5-flash?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-2.5-flash?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-2.5-flash for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.