gemini-3.1-flash-liteCompare gemini-3.1-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.
SentencePiece (Gemini)Scores on standardized evaluations. Higher percentages are better — and rank percentile shows
Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.1 Flash Lite
This model is not included in the current benchmark snapshot.
Missing models or measurements are not zero scores.
Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.
Gemini 3.1 Flash-Lite 适合输入多、任务边界清楚的应用。跨境客服可以用它翻译消息和评论;内容团队可以整理录音、提取视频要点;业务系统可以从 PDF 和图片中抽取字段,再生成简短摘要。它侧重低延迟与成本效率,便于把重复劳动放进自动化流程。
它支持文本、图片、音频、视频和 PDF 输入,输出文字。音频可以直接用于转写或摘要,文档可以用于提取实体、条款与待办信息。分析媒体时,最好明确是需要逐字记录、重点概括还是按条件筛选,避免把转写与总结混在一起。
结构化 JSON 输出适合评论分析、工单分类和表单整理。例如,对一条商品评价返回涉及属性、情绪、原文证据和退货意向,系统就能据此形成可统计的数据。先定义字段含义与合法取值,再用代表样本检查结果,可以减少不同批次之间的口径变化。
它也适合做任务路由:先判断请求是否明确、是否需要多步操作,再把复杂工作交给其他处理流程。函数调用可以连接业务工具,搜索和代码执行等能力需要在调用端启用。需要更多分析的任务可以提高思考强度;简单翻译则可以约束只返回译文,减少多余解释。
能省掉独立转写流程吗? 对支持的音频输入,可以直接要求输出文本;若要做实时字幕或严格时间对齐,需要单独确认应用使用的接口与输出要求。
怎样避免摘要遗漏? 规定必须覆盖的主题,例如决定、负责人和截止时间,并要求缺失项明确标记。
翻译请求说明目标语言、语气和术语,并要求只输出译文。转写请求则先说明要保留哪些口头细节。
请处理这段项目例会录音。
先按议题概括讨论,再提取决定和行动项。
行动项返回 JSON 数组,字段为 task、owner、due_date、evidence。
负责人和日期只有明确说出时才填写,缺失写 null;evidence 保留支持该任务的简短原话。
PDF 提取可以增加页码字段;评论分类应附类别定义和边界样例。用于任务路由时,要求只返回约定的路线值与简短原因,让程序容易接续处理。
/v1beta/models/gemini-3.1-flash-lite:generateContent| Parameter | Type | Default / range | Description |
|---|---|---|---|
temperature | number | = 10 ~ 2 | Sampling temperature; lower is more deterministic |
top_p | number | = 10 ~ 1 | Nucleus sampling probability mass |
max_tokens | integer | >= 1 | Maximum number of tokens in the response |
frequency_penalty | number | = 0-2 ~ 2 | Penalises repetition of frequent tokens |
presence_penalty | number | = 0-2 ~ 2 | Encourages introducing new topics |
stop | array | — | Up to 4 strings that stop generation |
seed | integer | — | Deterministic sampling seed (best-effort) |
n | integer | = 1>= 1 | Number of completions to generate |
stream | boolean | = false | Stream tokens via Server-Sent Events |
response_format | object | — | Force JSON object or schema-conforming output |
tools | array | — | Tool / function declarations the model may call |
tool_choice | string | autononerequired | Tool-choice policy or specific tool name |
logprobs | boolean | = false | Return per-token log probabilities |
top_logprobs | integer | 0 ~ 20 | Number of top log probabilities returned per token |
logit_bias | object | — | Per-token logit bias map |
user | string | — | End-user identifier for abuse monitoring |
Replace <YOUR_API_KEY> with the API key from your token settings.
All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.
Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.
| Parameter | Type | Default / range | Description |
|---|---|---|---|
temperature | number | = 10 ~ 2 | Sampling temperature; lower is more deterministic |
top_p | number | = 10 ~ 1 | Nucleus sampling probability mass |
max_tokens | integer | >= 1 | Maximum number of tokens in the response |
frequency_penalty | number | = 0-2 ~ 2 | Penalises repetition of frequent tokens |
presence_penalty | number | = 0-2 ~ 2 | Encourages introducing new topics |
stop | array | — | Up to 4 strings that stop generation |
seed | integer | — | Deterministic sampling seed (best-effort) |
n | integer | = 1>= 1 | Number of completions to generate |
stream | boolean | = false | Stream tokens via Server-Sent Events |
response_format | object | — | Force JSON object or schema-conforming output |
tools | array | — | Tool / function declarations the model may call |
tool_choice | string | autononerequired | Tool-choice policy or specific tool name |
logprobs | boolean | = false | Return per-token log probabilities |
top_logprobs | integer | 0 ~ 20 | Number of top log probabilities returned per token |
logit_bias | object | — | Per-token logit bias map |
user | string | — | End-user identifier for abuse monitoring |
| Supplier | RPM | TPM | RPD |
|---|---|---|---|
| default | Unlimited | Unlimited | Unlimited |
| Google AI Studio | Unlimited | Unlimited | Unlimited |
| Google Vertex | Unlimited | Unlimited | Unlimited |
No restriction
Compare gemini-3.1-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.
Create an API key with access to gemini-3.1-flash-lite, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.
Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.
The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.
The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.
Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.
