text-embedding-v4Compare text-embedding-v4 API pricing, supported endpoints, capabilities and access options on Modelsell.
BPE (vendor-specific)Scores on standardized evaluations. Higher percentages are better — and rank percentile shows
Metrics sourced fromArtificial Analysis 2026-09-30·Text Embedding V4
This model is not included in the current benchmark snapshot.
Missing models or measurements are not zero scores.
Benchmark charts preserve the source model selection and reasoning settings. Missing models or measurements are not zero scores, and benchmark cost or speed is not this site’s service commitment.
Text Embedding V4 把文本转换成数值向量,让应用能够比较内容含义,而不只依赖相同关键词。例如用户搜索“付款后多久到账”,知识库中的“充值到账时间”也可能与它相关。它适合企业文档检索、客服知识库、文章聚类、相似问题合并,以及为问答模型提供相关资料的 RAG 流程。
模型基于 Qwen3 训练,面向多语言文本表示。单条输入支持最多 8,192 Token,向量维度可在 64 至 2,048 之间选择,默认 1,024 维。较低维度便于控制存储和检索开销,较高维度可以作为追求检索效果时的候选;具体选择应通过自己的问题与文档测试。
返回的是向量,不是自然语言答案。典型流程是先将文档切成有意义的片段并建立向量索引,再把用户问题编码后检索相近内容,最后由业务程序或问答模型组织答案。需要更精细地筛选候选资料时,还可以在召回之后加入重排序。
尽量保留完整的段落、标题和必要上下文,避免在关键定义中间截断。过大的片段会混入多个主题,过小的片段又可能丢掉解释条件。对于订单规则、产品手册等资料,可用真实问题检查召回的片段是否足以支撑回答。
查询与文档应使用相同模型和兼容的维度、预处理方式。更换这些条件时,需要重新编码相应数据并建立匹配的索引,避免把不同空间中的向量直接比较。
通过 Embeddings 接口提交文本,并统一设置模型与向量维度。先把“退款申请通过后,款项会按原支付路径退回,到账时间以支付机构处理为准”等知识片段编码入库,再对“退款会退到哪里”生成查询向量,用相似度检索候选片段。
保存文档 ID、原文、模型名和维度,方便把检索结果还原为可阅读资料。用包含同义表达、简短问题和容易混淆问题的测试集检查召回,再调整切块与维度。返回向量后还需检索或分类程序处理,不能把数值数组直接当作用户答案。
Use the linked documentation for model-specific request parameters and examples.
All requests must include Authorization: Bearer <TOKEN> header. Anthropic-formatted endpoints accept the x-api-key header instead.
Generate tokens from the Tokens page; you can scope them to specific models, groups, IPs, and rate-limits.
| Supplier | RPM | TPM | RPD |
|---|---|---|---|
| Alibaba | Unlimited | Unlimited | Unlimited |
| official | Unlimited | Unlimited | Unlimited |
No restriction
Compare text-embedding-v4 API pricing, supported endpoints, capabilities and access options on Modelsell.
Create an API key with access to text-embedding-v4, then use the exact model ID and a supported endpoint from the API access section. Request fields depend on the selected endpoint.
Pricing depends on the selected provider group and the model billing unit. The current input, output, request, or media prices are shown on this page before sign-up.
The model catalog lists a context window of 8192 tokens. Check the selected endpoint for request limits.
Start with the use cases and prompting guidance on this page, then evaluate the model with representative inputs from your project.
