Gemini

Gemini 3.5 Flash Lite

GoogleToken-based
Alias:gemini-3.5-flash-lite
Create API Key

Compare gemini-3.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

textimageaudiovideofilecachingcode_interpreterfunction_callingtoolsweb_searchstructured_outputjson_modereasoningvisioncontext:1048576
Starting price
Input / Output · 1M
Context
1M
Maximum input window
Max output
65.5K
Maximum tokens per response
Modalities
→
Knowledge cutoff
Mar 2026
Released
Jul 2026

Pricing by Supplier

Google AI Studio
-50%
谷歌官方接口
Input$0.3$0.15/ 1M
Output$1.2$0.6/ 1M
Google Vertex
-30%
谷歌官方接口
Input$0.3$0.21/ 1M
Output$1.2$0.84/ 1M

Capabilities / Supported modalities

Prompt cachingCode interpreterFunction callingToolsWeb searchStructured outputReasoningVision
Input
Output

Provider & data privacy

Provider
GoogleDocs
Tokenizer
SentencePiece (Gemini)
License
Proprietary (commercial)Proprietary
Data retention41 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Gemini 3.5 Flash Lite

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Gemini 3.5 Flash Lite

为大量材料完成清楚、重复的处理

Gemini 3.5 Flash-Lite 是强调响应效率和成本控制的多模态模型,适合文档解析、内容分类、翻译和简单信息提取。它能接收文字、图片、音频、视频与 PDF,统一返回文本结果,方便把不同形式的材料送入同一业务流程。

例如处理培训资料时,可以从 PDF 提取课程主题,从录像整理章节摘要,再把多语言内容翻译成统一用语;处理业务文档时,则可以识别材料类型、提取关键字段并标记缺失项。任务数量多时,清楚的规则和稳定的输出结构比宽泛的要求更有帮助。

它也适合在智能体系统中承担范围明确的子任务,例如先筛选相关文档、归纳某一段材料或整理搜索结果,再交给后续任务综合判断。模型支持函数调用、结构化输出、搜索辅助和代码执行等能力,应用可以按需要提供工具,让提取与查询环节衔接起来。

多模态输入并不意味着直接生成对应媒体:它可以理解音视频,再用文字总结或回答问题。涉及小字、模糊录音或复杂页面时,应保留材料位置与无法判断的字段。先用一组实际样例确认分类边界和提取质量,再扩大处理范围,更容易发现规则中的遗漏。

Use cases and prompting

文档入库提示词

“请为每份培训材料生成入库记录。提取标题、主题、适用人群、前置知识和三条内容摘要;PDF 标注页码,视频标注相关片段。原文没有的信息填 null,专业术语沿用附件术语表,最后列出需要人工复查的材料。”

常见问题

  • 多种材料能按同一规则处理吗?可以先定义统一字段,再分别要求保留页码、时间位置或原文段落。
  • 怎样让翻译用词一致?提供术语表和目标读者,要求专名、缩写与计量单位按约定处理。
  • 能输出语音吗?这个模型返回文本;需要语音成品时,可将确认后的文字交给语音合成流程。

API access

API documentation

Code samples

RequestPOST/v1beta/models/gemini-3.5-flash-lite:generateContent
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
defaultUnlimitedUnlimitedUnlimited
Google AI StudioUnlimitedUnlimitedUnlimited
Google VertexUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gemini-3.5-flash-lite

What is gemini-3.5-flash-lite?

Compare gemini-3.5-flash-lite API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gemini-3.5-flash-lite?

Create an API key with access to gemini-3.5-flash-lite, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gemini-3.5-flash-lite priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gemini-3.5-flash-lite?

The model catalog lists a context window of 1048576 tokens. Check the selected endpoint for request limits.

What is the maximum output of gemini-3.5-flash-lite?

The model catalog lists a maximum output of 65536 tokens. Your request settings may set a lower limit.

How should I evaluate gemini-3.5-flash-lite for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.