gemini-3.5-flash-liteCompare gemini-3.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.
SentencePiece (Gemini)Scores on standardized evaluations. Higher percentages are better — and rank percentile shows
Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.5 Flash Lite
This model is not included in the current benchmark snapshot.
Missing models or measurements are not zero scores.
Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.
Gemini 3.5 Flash-Lite 是强调响应效率和成本控制的多模态模型,适合文档解析、内容分类、翻译和简单信息提取。它能接收文字、图片、音频、视频与 PDF,统一返回文本结果,方便把不同形式的材料送入同一业务流程。
例如处理培训资料时,可以从 PDF 提取课程主题,从录像整理章节摘要,再把多语言内容翻译成统一用语;处理业务文档时,则可以识别材料类型、提取关键字段并标记缺失项。任务数量多时,清楚的规则和稳定的输出结构比宽泛的要求更有帮助。
它也适合在智能体系统中承担范围明确的子任务,例如先筛选相关文档、归纳某一段材料或整理搜索结果,再交给后续任务综合判断。模型支持函数调用、结构化输出、搜索辅助和代码执行等能力,应用可以按需要提供工具,让提取与查询环节衔接起来。
多模态输入并不意味着直接生成对应媒体:它可以理解音视频,再用文字总结或回答问题。涉及小字、模糊录音或复杂页面时,应保留材料位置与无法判断的字段。先用一组实际样例确认分类边界和提取质量,再扩大处理范围,更容易发现规则中的遗漏。
“请为每份培训材料生成入库记录。提取标题、主题、适用人群、前置知识和三条内容摘要;PDF 标注页码,视频标注相关片段。原文没有的信息填 null,专业术语沿用附件术语表,最后列出需要人工复查的材料。”
/v1beta/models/gemini-3.5-flash-lite:generateContent| Parameter | Type | Default / range | Description |
|---|---|---|---|
temperature | number | = 10 ~ 2 | Sampling temperature; lower is more deterministic |
top_p | number | = 10 ~ 1 | Nucleus sampling probability mass |
max_tokens | integer | >= 1 | Maximum number of tokens in the response |
frequency_penalty | number | = 0-2 ~ 2 | Penalises repetition of frequent tokens |
presence_penalty | number | = 0-2 ~ 2 | Encourages introducing new topics |
stop | array | — | Up to 4 strings that stop generation |
seed | integer | — | Deterministic sampling seed (best-effort) |
n | integer | = 1>= 1 | Number of completions to generate |
stream | boolean | = false | Stream tokens via Server-Sent Events |
response_format | object | — | Force JSON object or schema-conforming output |
tools | array | — | Tool / function declarations the model may call |
tool_choice | string | autononerequired | Tool-choice policy or specific tool name |
logprobs | boolean | = false | Return per-token log probabilities |
top_logprobs | integer | 0 ~ 20 | Number of top log probabilities returned per token |
logit_bias | object | — | Per-token logit bias map |
user | string | — | End-user identifier for abuse monitoring |
Replace <YOUR_API_KEY> with the API key from your token settings.
All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.
Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.
| Parameter | Type | Default / range | Description |
|---|---|---|---|
temperature | number | = 10 ~ 2 | Sampling temperature; lower is more deterministic |
top_p | number | = 10 ~ 1 | Nucleus sampling probability mass |
max_tokens | integer | >= 1 | Maximum number of tokens in the response |
frequency_penalty | number | = 0-2 ~ 2 | Penalises repetition of frequent tokens |
presence_penalty | number | = 0-2 ~ 2 | Encourages introducing new topics |
stop | array | — | Up to 4 strings that stop generation |
seed | integer | — | Deterministic sampling seed (best-effort) |
n | integer | = 1>= 1 | Number of completions to generate |
stream | boolean | = false | Stream tokens via Server-Sent Events |
response_format | object | — | Force JSON object or schema-conforming output |
tools | array | — | Tool / function declarations the model may call |
tool_choice | string | autononerequired | Tool-choice policy or specific tool name |
logprobs | boolean | = false | Return per-token log probabilities |
top_logprobs | integer | 0 ~ 20 | Number of top log probabilities returned per token |
logit_bias | object | — | Per-token logit bias map |
user | string | — | End-user identifier for abuse monitoring |
| Supplier | RPM | TPM | RPD |
|---|---|---|---|
| default | Unlimited | Unlimited | Unlimited |
| Google AI Studio | Unlimited | Unlimited | Unlimited |
| Google Vertex | Unlimited | Unlimited | Unlimited |
No restriction
Compare gemini-3.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.
Create an API key with access to gemini-3.5-flash-lite, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.
Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.
The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.
The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.
Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.
