Gemini

Gemini 3.1 Flash Lite

GoogleToken-based
Alias:gemini-3.1-flash-lite
Create API Key

Compare gemini-3.1-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Released
May 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.25$0.125/ 1M
Output$1.5$0.75/ 1M
Cache Read$0.025$0.0125/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.25$0.175/ 1M
Output$1.5$1.05/ 1M
Cache Read$0.025$0.0175/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention78 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.1 Flash Lite

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.1 Flash Lite

为大量日常内容减轻处理负担

Gemini 3.1 Flash-Lite 适合输入多、任务边界清楚的应用。跨境客服可以用它翻译消息和评论;内容团队可以整理录音、提取视频要点;业务系统可以从 PDF 和图片中抽取字段,再生成简短摘要。它侧重低延迟与成本效率,便于把重复劳动放进自动化流程。

它支持文本、图片、音频、视频和 PDF 输入,输出文字。音频可以直接用于转写或摘要,文档可以用于提取实体、条款与待办信息。分析媒体时,最好明确是需要逐字记录、重点概括还是按条件筛选,避免把转写与总结混在一起。

结构化 JSON 输出适合评论分析、工单分类和表单整理。例如,对一条商品评价返回涉及属性、情绪、原文证据和退货意向,系统就能据此形成可统计的数据。先定义字段含义与合法取值,再用代表样本检查结果,可以减少不同批次之间的口径变化。

它也适合做任务路由:先判断请求是否明确、是否需要多步操作,再把复杂工作交给其他处理流程。函数调用可以连接业务工具,搜索和代码执行等能力需要在调用端启用。需要更多分析的任务可以提高思考强度;简单翻译则可以约束只返回译文,减少多余解释。

使用问题

能省掉独立转写流程吗? 对支持的音频输入,可以直接要求输出文本;若要做实时字幕或严格时间对齐,需要单独确认应用使用的接口与输出要求。

怎样避免摘要遗漏? 规定必须覆盖的主题,例如决定、负责人和截止时间,并要求缺失项明确标记。

Use cases and prompting

为每种重复任务固定输出规则

翻译请求说明目标语言、语气和术语,并要求只输出译文。转写请求则先说明要保留哪些口头细节。

请处理这段项目例会录音。
先按议题概括讨论,再提取决定和行动项。
行动项返回 JSON 数组,字段为 task、owner、due_date、evidence。
负责人和日期只有明确说出时才填写,缺失写 null;evidence 保留支持该任务的简短原话。

PDF 提取可以增加页码字段;评论分类应附类别定义和边界样例。用于任务路由时,要求只返回约定的路线值与简短原因,让程序容易接续处理。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.1-flash-lite:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.1-flash-lite

What is gemini-3.1-flash-lite?

Compare gemini-3.1-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.1-flash-lite?

Create an API key with access to gemini-3.1-flash-lite, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.1-flash-lite priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.1-flash-lite?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.1-flash-lite?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.1-flash-lite for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.