BaiLian

Text Embedding V4

AlibabaToken-based
Alias:text-embedding-v4
Create API Key

Compare text-embedding-v4 API pricing, supported endpoints, capabilities and access options on Modelsell.

embeddings语义检索RAGtext
Starting price
Input / Output · 1M
Context
8.2K
Maximum input window
Modalities
→
Released
Jun 2025

Pricing by Supplier

official
官方接口直连
Input$75/ 1M
Output$75/ 1M
Alibaba
阿里巴巴百炼官方
Input$75/ 1M
Output$75/ 1M

Capabilities / Supported modalities

Embeddings
Input
Output

Provider & data privacy

Provider
Unknown
Tokenizer
BPE (vendor-specific)
License
Provider-specificUnknown
Data retention59 daysNot used for upstream training by default

Performance

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced fromArtificial Analysis 2026-09-30·Text Embedding V4

This model is not included in the current benchmark snapshot.

Missing models or measurements are not zero scores.

Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.

About Text Embedding V4

用语义找到相关内容

Text Embedding V4 把文本转换成数值向量,让应用能够比较内容含义,而不只依赖相同关键词。例如用户搜索“付款后多久到账”,知识库中的“充值到账时间”也可能与它相关。它适合企业文档检索、客服知识库、文章聚类、相似问题合并,以及为问答模型提供相关资料的 RAG 流程。

模型基于 Qwen3 训练,面向多语言文本表示。单条输入支持最多 8,192 Token,向量维度可在 64 至 2,048 之间选择,默认 1,024 维。较低维度便于控制存储和检索开销,较高维度可以作为追求检索效果时的候选;具体选择应通过自己的问题与文档测试。

它会直接回答用户问题吗?

返回的是向量,不是自然语言答案。典型流程是先将文档切成有意义的片段并建立向量索引,再把用户问题编码后检索相近内容,最后由业务程序或问答模型组织答案。需要更精细地筛选候选资料时,还可以在召回之后加入重排序。

文档应该怎样切分?

尽量保留完整的段落、标题和必要上下文,避免在关键定义中间截断。过大的片段会混入多个主题,过小的片段又可能丢掉解释条件。对于订单规则、产品手册等资料,可用真实问题检查召回的片段是否足以支撑回答。

改维度后能沿用旧索引吗?

查询与文档应使用相同模型和兼容的维度、预处理方式。更换这些条件时,需要重新编码相应数据并建立匹配的索引,避免把不同空间中的向量直接比较。

Use cases and prompting

通过 Embeddings 接口提交文本,并统一设置模型与向量维度。先把“退款申请通过后,款项会按原支付路径退回,到账时间以支付机构处理为准”等知识片段编码入库,再对“退款会退到哪里”生成查询向量,用相似度检索候选片段。

保存文档 ID、原文、模型名和维度,方便把检索结果还原为可阅读资料。用包含同义表达、简短问题和容易混淆问题的测试集检查召回,再调整切块与维度。返回向量后还需检索或分类程序处理,不能把数值数组直接当作用户答案。

API access

API documentation

text-embedding-v4

Use the linked documentation for model-specific request parameters and examples.

Authentication

All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.

Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.

Rate limits

SupplierRPMTPMRPD
AlibabaUnlimitedUnlimitedUnlimited
officialUnlimitedUnlimitedUnlimited

No restriction

Frequently asked questions about text-embedding-v4

What is text-embedding-v4?

Compare text-embedding-v4 API pricing, supported endpoints, capabilities and access options on Modelsell.

How do I call text-embedding-v4?

Create an API key with access to text-embedding-v4, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.

How is text-embedding-v4 priced?

Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.

What is the context window of text-embedding-v4?

The model catalog lists a context window of 8192 tokens. Check the selected endpoint for request limits.

How should I evaluate text-embedding-v4 for my project?

Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.