Gemini

Gemini 2.5 Flash Lite

GoogleToken-based
Alias:gemini-2.5-flash-lite
Create API Key

Compare gemini-2.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Jan 2025
Released
Jul 2025

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.1$0.05/ 1M
Output$0.4$0.2/ 1M
Cache Read$0.01$0.005/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.1$0.07/ 1M
Output$0.4$0.28/ 1M
Cache Read$0.01$0.007/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention42 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 2.5 Flash Lite

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 2.5 Flash Lite

给重复内容安排轻量处理

Gemini 2.5 Flash-Lite 适合大量、简单且规则明确的工作,例如给消息打标签、提取单据字段、概括短文,或为上传内容生成简短说明。对于调用频繁、响应时间与预算敏感的应用,可以让它先完成基础整理,再把复杂问题交给后续流程。

它能接受文字、图片、音频、视频和 PDF,输出文字。你可以用一套字段规则处理不同形态的材料:从宣传截图提取活动时间,从文档中找联系人,从录音中概括主要诉求,再把结果返回给系统。请求越聚焦,越容易保持快速、稳定的处理方式。

结构化输出适合建立自动化管道。为每个字段定义含义、合法值和缺失规则,要求保留原始材料中的关键文字,程序便可以进行校验和统计。分类任务应包括容易混淆的类别说明;初筛任务则需要明确哪些情况进入人工复核,而不是强行得出确定判断。

它也支持思考与函数调用,但产品定位仍以轻量任务为主。需要检索最新资料、读取指定网页或执行计算时,应由调用端启用相应工具。涉及复杂推理、长链代码排错或重大方案取舍时,可以让它先提取事实和待解决问题,再交给更适合的深入分析流程。

应用设计问题

怎样减少批量提取的口径变化? 固定字段定义与示例,验证模糊输入、缺失内容和边界值,并让程序检查返回格式。

能处理图片里的小字吗? 可以尝试读取,但应提供清晰图片;关键号码与金额要求返回原文并复核,不要让它补齐模糊数字。

Use cases and prompting

用少量明确字段完成初筛

规定每次只做一类处理,避免把简单提取扩展成开放分析。

从这张活动海报中提取 title、start_date、end_date、location 和 signup_method。
仅返回 JSON。未写明的字段为 null,日期保留海报原文,不补充年份。
如果字迹不清或两处信息冲突,增加 needs_review: true,并在 note 中说明具体位置。

文本分类提供固定标签,短摘要指定字数与必含信息。先用代表样本检查字段准确率和复核比例,再接入批量处理;对复杂输入设置转交条件,使后续流程有明确入口。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-2.5-flash-lite:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-2.5-flash-lite

What is gemini-2.5-flash-lite?

Compare gemini-2.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-2.5-flash-lite?

Create an API key with access to gemini-2.5-flash-lite, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-2.5-flash-lite priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-2.5-flash-lite?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-2.5-flash-lite?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-2.5-flash-lite for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.