Skip to main content
Open SourceBAAI

BAAI bge-large-en-v1.5

BAAI bge-large-en-v1.5 is an English text embedding model from the Beijing Academy of Artificial Intelligence, used to turn text into dense vectors for retrieval and similarity search.

License Open Source
Contact providers for pricing

API Pricing

No pricing data available for this model at the moment.

Prices updated daily. Last check: Oct 5, 2026

Model Details

General

Creator
BAAI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
Yes
Aliases
BAAI/bge-large-en-v1.5

Strengths & Limitations

Strengths

  • Purpose-built embedding model — outputs dense vectors suited to semantic search, clustering, and classification rather than repurposing a chat model
  • Openly available weights, so the same model can be used via hosted APIs or self-hosted on your own GPUs without changing vector dimensions
  • Largest of the BGE English v1.5 size tiers, giving higher retrieval quality than the base and small variants
  • v1.5 revision improved similarity score distribution over the original BGE English release
  • Widely supported across inference providers, vector databases, and libraries such as Sentence-Transformers and LangChain, making it easy to swap in
  • Pairs directly with BAAI's bge-reranker models for a two-stage retrieve-then-rerank pipeline
  • Encoder architecture means low, predictable latency per request compared with generative models

Limitations

  • English-focused — not intended for multilingual corpora, where BAAI's multilingual BGE models or other options fit better
  • BERT-style encoder with a short maximum input length, so long documents must be chunked before embedding
  • Larger and slower than bge-base-en-v1.5 and bge-small-en-v1.5, with higher memory use if self-hosted
  • Retrieval quality still benefits from adding a query instruction prefix, an extra implementation detail that is easy to get wrong
  • Embeddings are not comparable with those from other models — changing embedding models requires reindexing your whole corpus
  • No generative or chat capability; it cannot answer questions on its own

Key Features

•Dense text embeddings for English semantic search and similarity
•Fixed-dimension output vectors suitable for cosine-similarity or dot-product search
•Query instruction prefix convention for asymmetric retrieval (query vs. passage)
•Largest tier of the BGE English v1.5 family (large / base / small)
•Open weights, usable via hosted inference APIs or self-hosted deployment
•Compatible with Sentence-Transformers, Hugging Face Transformers, and common vector databases
•Designed to be paired with BAAI bge-reranker models for second-stage ranking
•Batch embedding of document chunks for corpus indexing

About BAAI bge-large-en-v1.5

BAAI bge-large-en-v1.5 is an embedding model released by BAAI (Beijing Academy of Artificial Intelligence) as part of the BGE (BAAI General Embedding) family. It sits at the "large" size point of the English v1.5 line, above the base and small variants, and is served by inference providers as an embedding endpoint rather than a chat model — you send text and receive a fixed-length dense vector. The model encodes English text into dense embeddings that can be compared with cosine similarity or dot product for semantic search, clustering, deduplication, and classification. As a BERT-style encoder, it works on relatively short inputs — passages, paragraphs, and chunks rather than whole documents — so retrieval pipelines built on it typically chunk source material before indexing. The v1.5 revision of the BGE English models was published to improve similarity-score distribution and reduce the need for a query instruction prefix, though BAAI's documentation still recommends adding a short query instruction for retrieval tasks (queries prefixed, passages left plain). We do not track a context window, image modality, or tool-calling behavior for this model, since none of those apply to an embedding endpoint in the usual sense. In practice, bge-large-en-v1.5 is a common default for English RAG and vector-search stacks: it is openly available, widely mirrored across inference providers and self-hosting runtimes, and pairs naturally with a reranker (BAAI ships bge-reranker models for that second stage). Compared with the smaller BGE English variants it trades throughput and memory footprint for higher retrieval quality; compared with newer multilingual embedding models it is narrower in language coverage but simpler and cheaper to run.

Common Use Cases

bge-large-en-v1.5 is aimed at the indexing and retrieval layer of English-language systems: building the vector index for a RAG chatbot, powering site or documentation search, finding near-duplicate content, clustering support tickets or reviews, and generating features for downstream text classifiers. Because it is the large tier of the BGE English line, it suits workloads where retrieval accuracy matters more than raw embedding throughput — knowledge bases, legal or technical document search, and internal search over long-lived corpora that only need to be embedded once. High-volume, latency-critical, or cost-sensitive pipelines that embed on every request may prefer the base or small variants, and multilingual corpora are better served by a multilingual embedding model. A common pattern is to retrieve a broad candidate set with bge-large-en-v1.5 and then narrow it with a cross-encoder reranker before passing context to a generative model.

Frequently Asked Questions

How much does bge-large-en-v1.5 cost to use?

Pricing depends on which provider serves the model and how they bill — embedding endpoints are typically charged per million input tokens, and rates differ between providers and between shared and dedicated capacity. Because the weights are openly available, self-hosting on your own GPUs is also an option, in which case cost is driven by instance hours rather than tokens. See the pricing table on this page for current per-provider figures.

What is bge-large-en-v1.5 best used for?

Embedding English text chunks for semantic search and RAG retrieval, plus related vector tasks like clustering, deduplication, and similarity scoring. It is not a chat model and cannot generate answers — it produces vectors that a search index or a separate generative model consumes.

How does it differ from bge-base-en-v1.5 and bge-small-en-v1.5?

They are the same family and training recipe at different model sizes. The large variant generally gives the strongest retrieval quality of the three but is slower and more memory-hungry; base and small trade some accuracy for higher throughput and lower serving cost. All three require separate indexes — embeddings from one are not interchangeable with another.

Do I need to add an instruction prefix to my queries?

BAAI's guidance for the BGE English models is to prepend a short retrieval instruction to queries while leaving indexed passages unprefixed. The v1.5 release reduced how much this matters, but for asymmetric search (short query, long passage) it is still the recommended setup. Whatever you choose, apply it consistently at index time and query time.

Can it handle non-English text or long documents?

It is trained for English, so multilingual corpora are better matched to a multilingual embedding model. It is also a BERT-style encoder with a short maximum input length, so long documents should be split into chunks before embedding, with each chunk indexed separately.