Published 27 September 2026 · MTEB and RTEB leaderboard ranks, model cards and API prices checked on 26 September 2026
As of 26 September 2026, the best open-source embedding model for most teams is Qwen3-Embedding from Alibaba's Qwen team: it is Apache 2.0, reads 32K tokens, ships in 0.6B, 4B and 8B sizes, and the 8B model ranks 5th on MTEB English v2 and 4th on MTEB Multilingual v2 with 95% or more of those tasks unseen in training. The highest-scoring open model outright is Microsoft's harrier-oss-v1-27b (MIT licence, 1st on MTEB Multilingual v2), and on the retrieval-only RTEB board the top open model is NVIDIA's Nemotron-3-Embed-8B. For semantic search, pick from the retrieval column rather than the overall score and pair the embedder with keyword search. For RAG, shortlist four models from those boards and let a labelled sample of your own queries decide.
Every rank below comes from the MTEB leaderboard API on 26 September 2026 (MTEB English v2 with 41 tasks, MTEB Multilingual v2 with 131 tasks, and RTEB beta, which mixes open and private retrieval datasets), every dimension, context length and licence comes from the model's own card, and every price comes from the vendor's pricing page on the same day. The one piece of evidence that is ours is a retrieval study we ran across 50,000 documents and 5,000 queries, published as File Search vs Vector Search for RAG, and it changes how much weight the embedding model deserves in the first place.
Embedding model comparison table, 26 September 2026
Ranks are the leaderboard's own rank field on each board. "n/r" means the model has not been run on enough of that board's tasks to receive a complete score, which is common for new or API-only models and says nothing about quality in either direction.
| Model | Weights and licence | Params | Dimensions (truncation) | Max tokens | MTEB Eng v2 | MTEB Multi v2 | RTEB | Price per 1M tokens |
|---|---|---|---|---|---|---|---|---|
| Qwen3-Embedding-8B | Open, Apache 2.0 | 8B | 4096 (32 to 4096) | 32K | #5 | #4 | #15 | Self-host |
| Qwen3-Embedding-4B | Open, Apache 2.0 | 4B | 2560 (MRL) | 32K | #7 | #6 | n/r | Self-host |
| Qwen3-Embedding-0.6B | Open, Apache 2.0 | 0.6B | 1024 (32 to 1024) | 32K | #27 | #20 | n/r | Self-host |
| harrier-oss-v1-27b | Open, MIT | 27B | 5376 | 32,768 | n/r | #1 | n/r | Self-host |
| harrier-oss-v1-0.6b | Open, MIT | 0.6B | 1024 | 32,768 | n/r | #10 | n/r | Self-host |
| Nemotron-3-Embed-8B | Open, OpenMDW-1.1 (commercial use stated) | 8B | 4096 (sliceable) | 32,768 | n/r | n/r | #2 | Self-host |
| Octen-Embedding-8B | Open, Apache 2.0 | 8B | 4096 | 32,768 | n/r | #7 | #3 | Self-host |
| EmbeddingGemma-300M | Open, Gemma terms | 0.3B | 768 (512, 256, 128) | 2K | #34 | #31 | n/r | Self-host |
| granite-embedding-311m-multilingual-r2 | Open, Apache 2.0 | 311M | 768 (512 to 128) | 32,768 | n/r | #73 | n/r | Self-host |
| BGE-M3 | Open, MIT | 568M | 1024 dense, plus sparse and multi-vector | 8,192 | n/r | n/r | n/r | Self-host |
| jina-embeddings-v5-text-small | Open, CC BY-NC 4.0 (non-commercial) | 677M | 1024 (32 to 1024) | 32,768 | #15 | #14 | n/r | Self-host or Jina API |
| gemini-embedding-001 | Closed | n/a | 3072 (128 to 3072) | 2,048 | #9 | #5 | #13 | Not listed on pricing page today |
| gemini-embedding-2 (text, image, video, audio, PDF) | Closed | n/a | 3072 (128 to 3072) | 8,192 | n/r | n/r | n/r (preview #25) | $0.20 ($0.10 batch) |
| voyage-4-large | Closed | n/a | 1024 (256, 512, 2048) | 32,000 | n/r | n/r | #6 (#1 at 2,048 dims with evolved prompts) | $0.12 |
| voyage-4 | Closed | n/a | 1024 (256, 512, 2048) | 32,000 | n/r | n/r | #12 | $0.06 |
| voyage-4-lite | Closed | n/a | 1024 (256, 512, 2048) | 32,000 | n/r | n/r | #22 | $0.02 |
| text-embedding-3-large | Closed | n/a | 3072 (shortenable) | 8,192 | n/r | n/r | #39 | $0.13 |
| text-embedding-3-small | Closed | n/a | 1536 (shortenable) | 8,192 | n/r | n/r | #53 | $0.02 |
| all-MiniLM-L6-v2 (common default) | Open, Apache 2.0 | 23M | 384 | 256 | #144 | #150 | #82 | Self-host |
Two things in that table deserve a second look before anything else. MTEB's number one on English v2 today is a closed 8B model released in May 2026, and the two open models directly behind it were each trained on roughly half of the benchmark's tasks (the leaderboard reports a zero-shot share of 48% for both), which is why they are not the default recommendation here even though they outrank Qwen3 on that board. And the most talked-about hosted models of 2026, Voyage AI's voyage-4 family and Google's gemini-embedding-2, have no complete MTEB v2 score at all, so a table built only from MTEB averages silently drops them.
Best open-source embedding model in 2026
The default: Qwen3-Embedding
Qwen3-Embedding is the default recommendation because it is near the top of every board it has fully run on, under a licence your legal team will not argue about. The 8B model scores 0.7523 mean(task) on MTEB English v2 (5th) and 0.7058 on Multilingual v2 (4th); its retrieval column on Multilingual v2 is 0.7088. The leaderboard lists 95% of English v2 tasks and 99% of Multilingual v2 tasks as zero-shot for it, so the score is not a product of training on the test.
Qwen3 also has the practical properties that decide production fit: a 32K-token context window on every size, Matryoshka-style truncation from the full width down to 32 dimensions, instruction-aware queries (the card reports that task instructions typically add 1% to 5%), and a matching Qwen3-Reranker at 0.6B, 4B and 8B. Start with the 0.6B model if you are serving on modest hardware, move to 4B or 8B only if your own evaluation says the gain is worth the GPU.
Highest score: harrier-oss-v1
Microsoft's harrier-oss-v1 family, released in March 2026 under MIT, holds first place on MTEB Multilingual v2 with a 0.7427 mean(task) for the 27B model and the highest retrieval column on that board at 0.7827. It comes in three sizes (270M at 640 dimensions, 0.6B at 1,024, 27B at 5,376), all with 32,768-token inputs. The 0.6B model is the more interesting one for most teams: it ranks 10th on Multilingual v2 at 0.6901, ten places above Qwen3-Embedding-0.6B, at the serving cost of a sub-billion-parameter encoder. The 27B model is a quality ceiling, not a default; its weights alone are roughly 54 GB in bf16, before you index a single document.
Retrieval-only leaders: Nemotron-3-Embed and Octen
RTEB is the board to watch if all you care about is retrieval, because it scores models on legal, finance, healthcare and code retrieval tasks and keeps part of the data private so it cannot be trained on. NVIDIA's Nemotron-3-Embed-8B (released 16 July 2026) is the top open model there at 2nd, behind only a voyage-4-large configuration at 2,048 dimensions with evolved prompts, and its card states it is ready for commercial use under OpenMDW-1.1. Octen-Embedding-8B, an Apache 2.0 model derived from Qwen3-Embedding-8B, sits 3rd on RTEB and 7th on Multilingual v2. If your RAG corpus is contracts, filings or clinical notes, those two belong on the shortlist next to Qwen3.
Small models for CPU, edge and tight latency budgets
Under a billion parameters, four models are worth testing today. Qwen3-Embedding-0.6B and harrier-oss-v1-0.6b are both 1,024-dimension, 32K-context models; harrier ranks higher on Multilingual v2 (10th against 20th), Qwen3 is the only one of the two with an English v2 score (27th). Google's EmbeddingGemma-300M is smaller still at 300M parameters and 768 dimensions (truncatable to 512, 256 or 128) with a 2K input limit, and Google reports 69.67 on MTEB English and 61.15 on Multilingual. IBM's granite-embedding-311m-multilingual-r2, released in April 2026, is Apache 2.0 with a 32,768-token context, which is unusual at 311M parameters.
If you are still on all-MiniLM-L6-v2, the gap is large and measurable: its English v2 retrieval column is 0.4292 against 0.6183 for Qwen3-Embedding-0.6B. That model was the embedder in our own 50,000-document study, and we said at the time that a stronger embedder would very likely raise the vector numbers.
Licences that look open and are not
Open weights and open licence are different things, and the difference is where a CIO gets involved. Jina AI's jina-embeddings-v5-text-small ranks 14th on Multilingual v2 and is a strong sub-billion model, but it is CC BY-NC 4.0, and the card says commercial use requires contacting Jina. NVIDIA's llama-embed-nemotron-8b (3rd on Multilingual v2) is marked "for non-commercial/research use only", NV-Embed-v2 is CC BY-NC 4.0, and Tencent's KaLM-Embedding-Gemma3-12B (2nd on Multilingual v2) ships under a custom community licence. Two of the top five open models on Multilingual v2 therefore need a legal read before production use, which is the strongest practical argument for the Apache and MIT options above.
Hardware for self-hosting
A rough sizing rule from parameter counts: weights in bf16 take about 2 bytes per parameter, so an 8B embedder needs roughly 16 GB of GPU memory before activations and batching, a 4B model roughly 8 GB, and a 0.6B model roughly 1.2 GB. That is arithmetic, not a benchmark, and long 32K inputs add activation memory on top. For CPU serving, stay under a billion parameters and cap input length.
Which embedding model is best for semantic search?
For semantic search, the best embedding model is the one that ranks highest on retrieval tasks for your language and domain, which today means Qwen3-Embedding-8B or harrier-oss-v1 among open models and voyage-4-large or gemini-embedding-001 among hosted ones, run inside a hybrid search system rather than on its own. The overall MTEB English v2 score averages seven task types (retrieval, semantic similarity, classification, clustering, pair classification, reranking and summarization), and only one of those is search. Sort by the retrieval column or use RTEB.
Hybrid beats pure vector more often than the averages suggest
Our 50,000-document study is the reason this post does not stop at a model name. Across five datasets, vector search led keyword search on MRR@10 by 0.6320 to 0.5325. Remove the natural-language-to-code dataset and the ranking inverts: exact-match keyword search led the remaining four, 0.5931 to 0.5614. On science exam questions, keyword search scored 0.8145 against 0.6142 for vectors, because the query term (mitochondria, say) is the key rather than an approximation of a concept. The run used all-MiniLM-L6-v2 at 384 dimensions, Tantivy 0.22.0 and ChromaDB on an Apple M4 on 14 January 2026, a single run scored by exact document-ID matching with no LLM judge. A stronger embedder narrows the gap; it does not remove the class of query where the exact term matters.
That is why hybrid search is the production default for semantic search: a keyword index for identifiers, names and rare terms, an embedding index for paraphrase and intent, and a merge step. BGE-M3 is worth knowing here because a single MIT-licensed model emits dense, sparse and multi-vector representations, so one model can serve both sides of a hybrid index.
Latency and the embedding tax
Query latency is dominated by where the model runs. Meilisearch's published measurements from an AWS London instance put a local all-MiniLM-L6-v2 at about 10 ms per query and a local bge-large-en-v1.5 at about 60 ms, against about 460 ms for OpenAI's text-embedding-3-small and 750 ms for text-embedding-3-large over the network. If you run search-as-you-type, a local sub-billion model is usually the only option that fits.
Indexing is the other cost. In our study, building the keyword index over 50,000 documents took 2.11 seconds and embedding plus indexing took 161.6 seconds, a factor of 76.4 on the same machine. That tax matters when an agent needs to search a document it read seconds ago, and it is paid again in full every time you change embedding model.
Queries, prefixes and rerankers
Most modern embedders expect you to tell them whether an input is a query or a document, and skipping that step quietly costs accuracy. Qwen3 and harrier take a task instruction on the query side, nomic-embed-text-v2-moe requires search_query: and search_document: prefixes, and Voyage's API takes an input_type of query or document. Use each model's documented format, or your comparison is testing your preprocessing rather than the model.
Short queries against long documents (asymmetric search) are where embeddings are weakest, because a five-word query vector sits a long way from a thousand-word document vector. Chunking to paragraph or section size helps, and a reranker run over the top candidates (Qwen3-Reranker is the matching open option) is the usual next step once the embedder is fixed.
Multilingual, code and multimodal search
For multilingual or cross-lingual search, use MTEB Multilingual v2 rather than English v2, and then test the specific languages your users type, because an average over more than 250 languages can hide a weak one. For code search, MTEB has a dedicated Code benchmark, and vector search's biggest win in our own study was natural language to code, 0.9143 against 0.2901 for keyword. For images, PDFs, audio or video in one space, gemini-embedding-2 is now generally available with text at $0.20 per 1M tokens; our post on image embeddings versus vision models covers when to embed and when to reason.
How to choose an embedding model for RAG, step by step
Choosing an embedding model for RAG is a constraint filter followed by an experiment on your own data. Public leaderboards are the shortlist; your labelled queries are the decision.
An embedding model turns a piece of text into a fixed-length vector so that passages with similar meaning land close together. A RAG pipeline uses it twice: once at indexing time, to embed every chunk of your corpus, and again on every request, to embed the user's question and find the nearest chunks by cosine similarity. Those chunks are all the language model gets to read, so the embedder sets a ceiling on answer quality before generation even starts.
1. Write down the hard constraints first. Licence (commercial use or not), whether text may leave your infrastructure, the languages you must support, your latency budget per query, and the monthly spend you can tolerate. These remove more candidates than any benchmark will.
2. Shortlist four models. Take the top of the retrieval column on MTEB English v2 or Multilingual v2 and the top of RTEB, filter by your constraints, and include one small model as a baseline. A sensible shortlist today for an English commercial product is Qwen3-Embedding-4B, harrier-oss-v1-0.6b, Nemotron-3-Embed-8B and one hosted model such as voyage-4 or text-embedding-3-small.
3. Build a labelled eval set from your own queries. Somewhere between 100 and 200 real questions is enough to separate models, each with the ID of the chunk or document that answers it. Include exact identifiers, paraphrases, questions whose answer sits in a table, and a few with no answer at all. Keep the corpus and chunking fixed across every run.
4. Score retrieval, not answers. Use nDCG@10 (rewards putting the right chunk near the top), Recall@10 (did it appear at all) and MRR@10 (how high the first correct hit ranked). Answer quality depends on the generator and the prompt too, so measure it separately.
5. Measure latency, storage and price alongside quality. A million vectors at 1,024 dimensions take about 4.1 GB as float32, about 1 GB as int8 and about 128 MB as binary; the same million at 4,096 dimensions take about 16.4 GB as float32. Embedding a billion tokens costs $20 with text-embedding-3-small or voyage-4-lite, $60 with voyage-4, $130 with text-embedding-3-large and $200 with gemini-embedding-2 at standard rates. Pick the smallest model and the smallest dimension your eval tolerates.
6. Match context length to your chunks, not the other way round. A 32K-token window does not mean you should embed 32K-token chunks; one vector for a whole document blurs every fact in it. Most RAG pipelines retrieve best on paragraph- or section-sized chunks, which every model in the table above can read.
7. Plan for the next model. Store raw text and chunk IDs separately from vectors so a re-embed is a batch job rather than a migration.
A minimal version of steps 3 and 4 fits in a few lines with Sentence Transformers:
from sentence_transformers import SentenceTransformer
import numpy as np
# queries: list[str]; gold: list[int] (index of the answering chunk); chunks: list[str]
def evaluate(model_id, queries, gold, chunks, k=10, **query_kwargs):
model = SentenceTransformer(model_id)
doc_vecs = model.encode(chunks, normalize_embeddings=True, batch_size=32)
q_vecs = model.encode(queries, normalize_embeddings=True, **query_kwargs)
ranks = np.argsort(-(q_vecs @ doc_vecs.T), axis=1)[:, :k]
hits = [np.where(r == g)[0] for r, g in zip(ranks, gold)]
recall = np.mean([len(h) > 0 for h in hits])
mrr = np.mean([1.0 / (h[0] + 1) if len(h) else 0.0 for h in hits])
return {"recall@10": recall, "mrr@10": mrr}
print(evaluate("Qwen/Qwen3-Embedding-0.6B", queries, gold, chunks, prompt_name="query"))
Run the same function for each shortlisted model with its documented query prompt, and keep the numbers next to latency and cost in one table.
When to fine-tune or use a domain model. If your text is full of vocabulary a general model rarely saw (drug names, statute citations, internal product codes), check for a domain model first and fine-tune a general one second. Results do move by domain: in a September 2026 AIMultiple benchmark of 11 open models, BGE Large EN v1.5 led the finance subset while ranking fourth overall.
Why the top of the leaderboard is often the wrong pick
Three problems sit underneath any single leaderboard number, and the 26 September snapshot shows all three. Training overlap: the leaderboard's own zero-shot field shows some top English v2 models trained on about half the tasks they are scored on. Coverage: several of 2026's most-used hosted models have no complete MTEB v2 score, so an average-sorted table cannot rank them. Task mix: the overall MTEB score blends clustering and classification with retrieval, and a model can climb the overall board without getting better at search.
Leaderboards also move. Of the models in the comparison table, harrier-oss-v1, jina-embeddings-v5, voyage-4, gemini-embedding-2, granite-embedding R2 and Nemotron-3-Embed were all released in 2026. Re-running your own eval set each quarter, or whenever a candidate appears near the top of RTEB, is cheaper than reacting to every release.
Switching costs and where your data goes
Changing embedding model means re-embedding everything, because vectors from different models are not comparable. Google's documentation says so directly for its own models: gemini-embedding-001 and gemini-embedding-2 produce incompatible spaces and upgrading requires re-embedding all existing data. Voyage is the exception worth knowing, because its documentation states that all voyage-4 series embeddings are compatible with each other, so you can mix voyage-4-lite queries with voyage-4-large documents.
Data residency is the other half of the decision. A hosted embedding API receives every chunk of your corpus and every user query; providers differ on retention and training use, and Google's pricing page lists free-tier inputs as used to improve its products while paid-tier inputs are not. Self-hosting an open model keeps text inside your own infrastructure at the price of GPU serving, upgrades and on-call. That trade is the one engineering managers should price explicitly, because it is recurring work rather than a one-time choice.
What a better embedding model will not fix
An embedding model measures similarity, and several common retrieval failures are not similarity problems, a list Galileo documents in detail. A query about "2024 Q3 revenue" retrieves revenue passages from every quarter, because the year carries almost no weight in the vector. A query that excludes something ("Python examples without external dependencies") retrieves the excluded thing. And in our study the multi-hop HotpotQA questions failed in a specific way: vector search found the document about the first entity and missed the bridge document about the second, because the bridge document does not look like the query. No embedder upgrade fixes that; it needs metadata filters, reranking, or a stored relationship.
Agent memory adds a harder version of the same problem. When the corpus is conversation history, it grows every turn and contradicts itself: the fact stated in March and the correction in June are both similar to the query, and the vector has no view on which is current, or on whether "my manager" and "Sarah" are the same person. Choosing a better embedding model for that corpus improves which passages come back; it does not decide which of them is still true.
Where Maximem Synap fits
Maximem Synap is a memory layer for AI agents, and one of the things it takes off your plate is this entire decision. Synap runs embedding on Maximem's own infrastructure, so text is not sent to a third-party embedding API for that step, and choosing and upgrading the embedding model is Synap's job rather than yours.
What sits around the embedder matters more. Synap reads a conversation and stores the facts, preferences and episodes worth keeping rather than raw transcript chunks; retrieval runs several techniques at once and merges them into one ranked answer (Resilient Retrieval); Entity Resolution collapses "my manager", "Sarah" and "Sarah Chen" into one person; and Conscious & Lossless Forgetting stops the agent repeating a fact that has since changed. In-conversation reads come from a cache pre-fetched into your own process, at an asserted P75 under 15 ms. If you are building document RAG, the steps above are the right way to pick a model. If you are building an agent that has to remember users, the embedding model is one component of a memory architecture, and Synap handles that architecture for you.
Frequently asked questions
What is the best open-source embedding model in 2026?
As of 26 September 2026, Qwen3-Embedding (0.6B, 4B or 8B, Apache 2.0, 32K context) is the best default open-source embedding model: the 8B model ranks 5th on MTEB English v2 and 4th on MTEB Multilingual v2, with 95% or more of tasks zero-shot. Microsoft's harrier-oss-v1-27b (MIT) has the highest MTEB Multilingual v2 score at 0.7427, and NVIDIA's Nemotron-3-Embed-8B is the top open model on the retrieval-only RTEB board.
Which embedding model is best for semantic search?
For semantic search, choose by the retrieval column or the RTEB board rather than the overall MTEB average. Among open models that points to Qwen3-Embedding-8B, harrier-oss-v1 and Nemotron-3-Embed-8B; among hosted ones, voyage-4-large (1st on RTEB) and gemini-embedding-001. Run whichever you pick inside hybrid search, because exact-match keyword search still wins on identifiers and rare terms; in Maximem's 50,000-document study keyword search won outright on the two entity-heavy datasets and led the average once code search was excluded.
How do I choose an embedding model for RAG?
List your hard constraints (licence, data residency, languages, latency, budget), shortlist four models from the retrieval column of MTEB and from RTEB, then score each on 100 to 200 of your own labelled queries with nDCG@10, Recall@10 and MRR@10 while also recording latency, storage and price. Pick the smallest model and dimension that meets your quality bar, and store raw text separately so a future re-embed is a batch job.
Is a bigger embedding model always better for RAG?
No. Bigger embedding models usually score higher on public boards, but the gain on your own data can be small while serving cost grows roughly with parameter count; an 8B model needs about 16 GB of GPU memory for weights alone, against about 1.2 GB for a 0.6B model. harrier-oss-v1-0.6b ranks 10th on MTEB Multilingual v2, within 0.06 of the 27B model's mean score at roughly one forty-fifth of the parameters.
Can I use Jina or NVIDIA embedding models commercially?
It depends on the exact model. jina-embeddings-v5-text-small and NV-Embed-v2 are CC BY-NC 4.0 and llama-embed-nemotron-8b is marked non-commercial or research use only, so all of those need a commercial agreement. NVIDIA's Nemotron-3-Embed-8B is released under OpenMDW-1.1 and its card states it is ready for commercial use. Always read the licence of the specific checkpoint rather than the model family.
Do I need to re-embed my data if I change embedding models?
Almost always. Vectors from different models live in different spaces and cannot be compared; Google states that gemini-embedding-001 and gemini-embedding-2 are incompatible and require re-embedding. Voyage's voyage-4 series is an exception within itself, since its documentation says all voyage-4 embeddings are mutually compatible. Budget the re-embed as compute time: in Maximem's study, embedding 50,000 documents locally took 161.6 seconds against 2.11 seconds for a keyword index.
Sources, retrieved 26 September 2026: MTEB leaderboard and its API (MTEB English v2, MTEB Multilingual v2, RTEB beta); model cards for Qwen3-Embedding-8B, Qwen3-Embedding-0.6B, harrier-oss-v1-27b, Nemotron-3-Embed-8B, llama-embed-nemotron-8b, Octen-Embedding-8B, EmbeddingGemma, granite-embedding-311m-multilingual-r2, BGE-M3, nomic-embed-text-v2-moe, jina-embeddings-v5-text-small, KaLM-Embedding-Gemma3-12B-2511; OpenAI embeddings guide and model pages for text-embedding-3-large and text-embedding-3-small; Gemini embeddings documentation and Gemini API pricing; Voyage AI embeddings and pricing; Meilisearch semantic search model comparison (updated 4 August 2026); AIMultiple open-source embedding benchmark (updated 25 September 2026); Maximem, File Search vs Vector Search for RAG.



