OpenAI

GPT 6 Luna

OpenAIToken-based
Alias:gpt-6-luna
Create API Key

Compare gpt-6-luna API pricing, supported endpoints, capabilities and access options on Modelsell.

textimagestreamingfunction_callingtoolsstructured_outputjson_modevisionweb_searchcachingreasoningcode_interpretercontext:1050000new
Starting price
Input / Output · 1M
Context
1.1M
Maximum input window
Max output
128K
Maximum tokens per response
Modalities
→
Knowledge cutoff
May 2026
Released
Sep 2026

Pricing by Supplier

OpenAI
-40%
Openai官方接口
Input$0.1$0.06/ 1M
Output$0.5$0.3/ 1M
Cache Read$0.01$0.006/ 1M
Cache Write (5m)$0.125$0.075/ 1M
Cache Write (1h)$0.2$0.12/ 1M
Azure
-40%
微软云API直连,稳定可靠
Input$0.1$0.06/ 1M
Output$0.5$0.3/ 1M
Cache Read$0.01$0.006/ 1M
Cache Write (5m)$0.125$0.075/ 1M
Cache Write (1h)$0.2$0.12/ 1M
GPT官+AZ 混合
-40%
适合生产环境,az 和官 key 混合渠道
Input$0.1$0.06/ 1M
Output$0.5$0.3/ 1M
Cache Read$0.01$0.006/ 1M
Cache Write (5m)$0.125$0.075/ 1M
Cache Write (1h)$0.2$0.12/ 1M
GPT特价
-50%
适合vibe coding自用,gpt pro 号池
Input$0.1$0.05/ 1M
Output$0.5$0.25/ 1M
Cache Read$0.01$0.005/ 1M
Cache Write (5m)$0.125$0.0625/ 1M
Cache Write (1h)$0.2$0.1/ 1M

Capabilities / Supported modalities

StreamingFunction callingToolsStructured outputVisionWeb searchPrompt cachingReasoningCode interpreter
Input
Output

Provider & data privacy

Provider
OpenAIDocs
Tokenizer
o200k_base
License
Proprietary (commercial)Proprietary
Data retention30 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·GPT 6 Luna

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About GPT 6 Luna

让重复工作更高效

GPT-6 Luna 面向范围明确、需要频繁执行的任务。它适合把客服留言分配到正确类别,从文档中提取固定字段,为商品资料生成简洁说明,或把零散记录整理成统一格式。对这类工作,清楚的规则和稳定的输出往往比长篇分析更重要。

文本与图片都能成为输入

模型支持文本和图片输入、文本输出。你可以让它读取截图中的关键信息,对照规则判断某份材料属于哪一类,或结合商品图和已提供的属性撰写说明。需要结构化结果时,提前定义字段、可选值和缺失信息的处理方式,便于应用直接接收和检查。

把规则写成可重复执行的流程

例如工单分流,应说明类别边界,提供容易混淆的例子,并为无法确定的内容设置人工复核选项。信息提取任务则要明确只能使用输入中出现的事实,不能为凑齐字段而补全不存在的信息。先用具有代表性的样本检查规则,再扩大使用范围,更容易发现边界问题。

Luna 也可参与工具工作流。内置工具和带推理的函数调用适合通过 Responses API 接入;Chat Completions 的函数调用要求 reasoning_effort=none。对于简单分类可采用较轻的推理设置,遇到复杂或高歧义材料时,再转交更深入的分析流程。

Use cases and prompting

客服分流示例

“把下面的留言分为退款、物流、产品咨询、其他四类。输出 category、summary、needs_human 三个字段。用户同时提到多个问题时,按最需要处理的问题分类;信息不足或存在投诉升级迹象时,needs_human 为 true。不要替用户补写未提到的事实。”

提供各类别的边界和几条真实样本,有助于形成一致结果。

信息缺失时该怎么输出?

在规则中指定 null、空数组或待确认状态,不要让模型猜测。程序端继续校验字段和取值范围。

哪些任务适合升级处理?

涉及多份矛盾材料、复杂因果判断或难以定义的需求时,先让Luna完成整理与分流,再交给更深入的分析环节。

API access

API documentation

Code samples

RequestPOST/v1/chat/completions
Example request
Parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Replace <YOUR_API_KEY> with the API key from your token settings.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Supported parameters

Generation parameters
ParameterTypeDefault / rangeDescription
temperature
number
=10 ~ 2
Sampling temperature; lower is more deterministic
top_p
number
=10 ~ 1
Nucleus sampling probability mass
max_tokens
integer>= 1Maximum number of tokens in the response
frequency_penalty
number
=0-2 ~ 2
Penalises repetition of frequent tokens
presence_penalty
number
=0-2 ~ 2
Encourages introducing new topics
stop
array—Up to 4 strings that stop generation
seed
integer—Deterministic sampling seed (best-effort)
n
integer
=1>= 1
Number of completions to generate
stream
boolean
=false
Stream tokens via Server-Sent Events
response_format
object—Force JSON object or schema-conforming output
tools
array—Tool / function declarations the model may call
tool_choice
string
autononerequired
Tool-choice policy or specific tool name
logprobs
boolean
=false
Return per-token log probabilities
top_logprobs
integer0 ~ 20Number of top log probabilities returned per token
logit_bias
object—Per-token logit bias map
user
string—End-user identifier for abuse monitoring

Rate limits

SupplierRPMTPMRPD
GPT特价UnlimitedUnlimitedUnlimited
OpenAIUnlimitedUnlimitedUnlimited
AzureUnlimitedUnlimitedUnlimited
GPT官+AZ 混合UnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about gpt-6-luna

What is gpt-6-luna?

Compare gpt-6-luna API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call gpt-6-luna?

Create an API key with access to gpt-6-luna, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is gpt-6-luna priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of gpt-6-luna?

The model catalog lists a context window of 1050000 tokens. Check the selected endpoint for request limits.

What is the maximum output of gpt-6-luna?

The model catalog lists a maximum output of 128000 tokens. Your request settings may set a lower limit.

How should I evaluate gpt-6-luna for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.