Back to all writing

Similar Text Isn’t Always the Right Answer: Choosing an Embedding Model

Choose embedding models for RAG and semantic search. Use practical shortlists to test result quality, response times and cost on your own data.

Blue glass query and document representations pass through a precision retrieval and reranking apparatus to an amber answer-bearing document

Someone types a question into your support application: “How do I reset my password?”

Your search returns: “I forgot my password. How can I change it?”

That match helps you find duplicate support questions. A person trying to reset their password needs instructions instead.

For that, a useful result might be: “Open Settings, select Security, and choose Reset password.”

Same query. Different correct result.

Before comparing embedding models, decide what a good match means for your application. Are you looking for similar meaning, a passage that answers a question, code that performs a particular task, or a product someone might want next?

Embeddings can help with all of these. What counts as a useful result changes with the task.

What an embedding actually gives you

An embedding represents an input as numbers. A typical dense text model produces a fixed-length list of numbers, called a vector. Comparing vectors lets you find items the model considers closely related.

What counts as “closely related” depends on the model's training data, training objectives and input instructions. A vector database finds nearby vectors, but that closeness still needs to mean something useful for your product.

In retrieval-augmented generation, or RAG, the application retrieves relevant evidence, then a separate language model uses it to write a response. The embedding model helps find evidence; it does not write the answer. Embeddings also work for ordinary search that returns results without generating a response.

Start by writing one sentence:

“Given ___, retrieve ___, because ___.”

For a support assistant: “Given a customer question, retrieve current instructions that resolve it, because the answer needs to be supported by our documentation.”

For duplicate detection: “Given a new ticket, retrieve tickets describing the same issue, because we want to combine repeated reports.”

That gives you a way to judge a model beyond its position on a leaderboard.

Q&A retrieval and similarity ask different questions

In symmetric matching, you compare inputs that play roughly the same role: a question with another question, or an article with another article. You are often asking whether they express similar meaning.

In asymmetric retrieval, the inputs play different roles: a question and a passage containing its answer, or a task description and a function that performs it. They can look quite different while still being a useful match. Sentence Transformers explains this distinction in its semantic-search guidance. Semantic-search documentation

The password example makes the difference concrete:

What the application needsWhat a good match looks like
Find another version of the question“I forgot my password. How can I change it?”
Find evidence for an answer“Open Settings, select Security, and choose Reset password.”
Find a particular failureA document mentioning the exact code ERR_AUTH_104

These examples show what we want the search to return; they are not results from a model benchmark.

A FAQ application can use either approach. Matching an incoming question to a stored FAQ question, then returning its attached answer, can work. Searching the answers themselves or a document collection needs a different check: do the returned passages answer the question? Calling the interface “Q&A” does not settle the choice.

The same model may support both tasks, sometimes with different instructions. You do not necessarily need a separate model family for each.

Here, “asymmetric” describes the task and how the inputs are encoded. Cosine similarity, a common way to compare vectors, is still symmetric: swapping two already-computed vectors does not change their score.

The input format belongs to the model

Some models need different query and document instructions. Others accept an input-role parameter or use the same encoding process for both.

For example, Voyage provides query and document input roles. Qwen3-Embedding's model card adds a prompt to the query and encodes the retrieval documents without that prompt. Follow the format documented for your model and task. Voyage input types, Qwen3 model card

An instruction or prefix from another model's tutorial may not help. A successful API response tells you that the request was accepted, not that the model received the right instructions for your task.

Bi-encoders retrieve; cross-encoders examine the pair

Symmetric and asymmetric describe what you want to match. Bi-encoder and cross-encoder describe how the model processes the inputs. A bi-encoder can support either matching task, depending on its training and input format.

A bi-encoder creates each input's embedding independently. You can calculate document embeddings in advance and store them in a search index. When a query arrives, embed it and search those document vectors. Reusing document embeddings makes searching large collections practical. The encoders may share learned parameters, or weights; “bi” does not require two separately trained models. Bi-encoder documentation

A cross-encoder reads the query and a candidate document together, then scores their relevance. It can notice distinctions the first search missed, but needs to score each pair. It generally does not produce reusable document vectors for a vector index, making it better suited to examining a shortlist than searching a large collection. Retrieve-and-rerank documentation

Reranking means scoring the retrieved candidates again to put the most relevant ones first. For a documentation assistant, a common flow is retrieve → rerank → answer:

  1. Retrieve candidate passages using embeddings, optionally combined with lexical, or keyword-based, search.
  2. Rerank those passages against the question.
  3. Give the selected evidence to the language model so it can write an answer.

The embedding model and reranker help select evidence. Neither writes the answer, and a reranker cannot recover a passage missing from its shortlist.

Strong examples worth comparing

These candidates were checked in October 2026. Which ones are worth testing depends on the quality you need and how you plan to run them. Open-weight models can be downloaded and run on your infrastructure; hosted models are accessed through an API operated by the provider.

RoleCandidateWhen to consider it
Open-weight bi-encoderQwen3-Embedding-0.6B, 4B and 8BMultilingual retrieval with task instructions. Start with 0.6B if compute is limited; compare the larger models if better results would justify the extra resources
Open-weight bi-encoder, using its dense outputBAAI/bge-m3Multilingual retrieval, with dense, sparse and multivector options. Its dense embeddings can be reused in a vector index; the other modes need different scoring and indexing approaches
Open-weight cross-encoderQwen3-Reranker-0.6B, 4B and 8BMultilingual pair scoring with task instructions. These use a decoder backbone to score the query and document together. Compare model sizes against quality and response time
Open-weight cross-encoderBAAI/bge-reranker-v2-m3An encoder-based multilingual reranker to test when serving efficiency matters. It returns pair scores; the similarly named bge-m3 returns embeddings
Hosted cross-encoderVoyage rerank-3 and rerank-3-liteReranking through an API. Voyage positions rerank-3 for quality and the lite version for speed and cost. Compare them using your documents and shortlist size

For a lightweight English passage-ranking baseline, or starting point for comparison, consider cross-encoder/ms-marco-MiniLM-L6-v2. It has roughly 23 million parameters.

If you prefer a hosted embedding model, the OpenAI models and Voyage models in the shortlist below can produce your query and document embeddings without requiring you to run an embedding server.

Reranking adds processing time and cost. Passage count and length, model size, request batching and hardware all matter. Measure the improvement alongside p95 latency: the response time within which 95% of measured requests finish. Average speed can hide slow requests your users experience.

A larger cross-encoder or a longer shortlist should improve the results enough to justify that extra work.

Choose for the job you are actually doing

Your application gives you a starting point and a way to test whether the model helps.

Documentation search and RAG: Start with models trained for retrieval. Check whether the returned passages actually support an answer, including any conditions or exceptions. A paragraph saying password resets are supported is less useful than the instructions for doing one.

Similarity and duplicate detection: Compare inputs with the same role, such as two tickets. Include small differences that matter: “refund approved” versus “refund rejected,” or similar incidents in different customer accounts. Two items can discuss the same topic without describing the same issue. If an incorrect merge would be costly, use embeddings to suggest candidates and add a stricter check before merging them.

Clustering and classification: Embeddings can help group related tickets or assign them to categories. Check whether those groups and labels are useful. Grouping tickets by customer name does not help a team that needs categories such as billing, login and delivery. For classification, test the full classifier or label-matching method against labelled examples. Good search results do not prove that tickets will be routed correctly.

Code search: “Find the function that retries failed requests” asks for code matching an intent. “Find every use of retry_failed_request” asks for an exact symbol. Include function signatures and useful surrounding context when testing embeddings, and compare keyword-based search too. A code-focused embedding model is worth testing for collections dominated by code.

Multilingual search: Test your users' languages, including mixed-language queries and searching documents in another language. A Hindi query finding an English manual is a different test from an English query finding it. A supported-language list does not establish quality for either case.

Multimodal search: This involves more than one type of content, such as text and images. If users search charts, scanned pages, photos, audio or video, check that the model can match the input types you need. Extracting text is a useful starting point, but it can lose important information. Embedding a chart's title will not help you find values that were never passed to the model.

Recommendations: Item-content embeddings can support “more like this.” Personalised recommendations also need the person's preferences, behaviour, constraints or feedback. Similar articles are not necessarily articles someone wants to read one after the other.

Dense, sparse and multivector are another decision

“Multilingual,” “dense” and “good for retrieval” can describe the same model. These terms answer different questions: which languages it handles, how it represents content, and what task it is suited to.

Dense embeddings commonly use one compact vector per chunk: a passage or other piece of content. They are a straightforward starting point when matching meaning is useful and your search system supports vector retrieval.

Learned sparse embeddings, such as SPLADE-style representations, use vectors with mostly zero values. Their non-zero values give weights to vocabulary terms, allowing the model to represent which terms matter for matching. They can complement dense search. Sparse-encoder documentation

Multivector models keep several vectors per item. ColBERT encodes the query and document independently, then compares their finer-grained representations through late interaction. Keeping that detail changes storage, indexing and scoring requirements. An index designed for one vector per item cannot automatically handle this approach. ColBERT paper

Include BM25 in the comparison too. It is a keyword-based ranking method rather than a neural embedding model. It gives you a useful reference point, particularly for exact identifiers, names and specialised terms.

Hybrid search combines results or scores from different retrieval methods, often keyword-based and dense search. Reranking happens after retrieval: it scores the candidates again, usually by examining each one alongside the query. You can use either or both, but they change different parts of the search process. Retrieve-and-rerank documentation

Add these steps when you find a problem they could solve. If you change the model, passage boundaries, method of combining results and reranker at once, it becomes difficult to tell which change helped.

A small model shortlist

Here are starting points for a few common requirements, with model details checked in October 2026. Use the shortlist to decide what to test, then compare performance on your own data.

Starting requirementCandidateWhy it belongs in the comparison
A hosted text baselineOpenAI text-embedding-3-smallA general text embedding API. Compare text-embedding-3-large if better results on your data would justify its extra cost
Hosted retrieval, including multilingual textVoyage 4 familyRetrieval-focused options. Voyage positions voyage-4-lite for speed and cost; test voyage-code-4 separately for code-heavy collections
Text and visually rich documentsCohere Embed 5Pro and Fast variants support text, images and mixed inputs
Search across text, images, audio and videoGemini Embedding 2Supports those modalities in a shared embedding space
Self-hosted multilingual textQwen3-Embedding-0.6BAn open-weight starting point that accepts task instructions. Consider it when you want to control deployment; larger versions are available

For sentence similarity or clustering on modest English datasets, the older all-MiniLM-L6-v2 is another lightweight local baseline. It produces 384-dimensional vectors. By default, it truncates inputs beyond 256 wordpieces, which are the model's text units and are not necessarily whole words. That limit makes it a poor default for embedding entire manuals. Model card

A model that meets your quality requirements and fits your existing setup may be a better choice than a slightly stronger one that requires another service to run and maintain.

The comparison that is worth running

I would start with a few dozen cases I can inspect by hand, then expand to hundreds as I discover where the models struggle. Keep a held-out set: examples you do not use for tuning, so you can check the final choice on fresh cases. A small evaluation helps you learn, but cannot guarantee performance across every situation.

1. Decide what counts as a correct result

For retrieval, identify which passages actually support an answer. Include partial matches, plausible but wrong answers, outdated documents, exact identifiers and questions your collection cannot answer.

For duplicates, label whether the items should be merged. For recommendations, decide what makes a suggestion useful. Have people review ambiguous cases before treating small differences in scores as meaningful.

When you split the examples into tuning and held-out sets, keep related examples together. Nearly identical documents or tickets from one incident appearing in both sets can make the system look better at handling new cases than it really is.

2. Keep the first comparison controlled

Compare models using the same document collection, labelled relevant passages, chunk boundaries and number of returned results. Configure each model according to its own documentation. Start by comparing keyword-based search with a few dense embedding models, before adding hybrid search or reranking.

For a small collection, compare the query against every stored vector using exact vector search. This separates embedding quality from approximation. Larger collections often use approximate nearest-neighbour (ANN) search for speed, which may miss nearby vectors. Test your deployed index too: its settings can change the results.

3. Measure the result you need

For retrieval, k is the number of top results you inspect. Three useful measures answer different questions:

  • Recall@k: How many of the known relevant items appear in those results, as a fraction of all the relevant items?
  • Hit rate@k: Do those results contain at least one relevant item?
  • nDCG@k: Are the most useful results near the top? This measure considers ranking position and can distinguish strongly relevant results from partially relevant ones.

Sentence Transformers provides evaluators for retrieval, similarity and classification tasks. Evaluation documentation

Here is a small function for recall and hit rate. Each ranked ID should be unique and refer to the same kind of item used in your labels, such as a passage:

def retrieval_metrics(ranked_ids, relevant_ids, k=5):
    if k < 1:
        raise ValueError("k must be positive")

    relevant = set(relevant_ids)
    if not relevant:
        # Evaluate unanswerable queries separately.
        return None

    retrieved = set(ranked_ids[:k])
    found = retrieved & relevant
    return {
        f"recall@{k}": len(found) / len(relevant),
        f"hit@{k}": float(bool(found)),
    }

For example, suppose the relevant passages are C and D, and the first three results are A, B and C. The function returns recall@3 = 0.5, because it found one of the two relevant passages, and hit@3 = 1.0, because it found at least one. A hit does not tell you that all the evidence you need was retrieved.

Average each metric across questions your collection can answer. Check unanswerable questions separately: does the full application avoid giving an unsupported answer? A top-k vector search can return nearby items even when none answers the question.

For duplicate detection, measure precision and recall at your chosen threshold. Precision tells you how many pairs flagged as duplicates really are duplicates; recall tells you how many true duplicates you found. A cosine score of 0.8 does not mean there is an 80% chance that two items are duplicates. Choose the threshold using validation examples, then check it on the held-out set.

Also measure p95 search latency, how quickly you can index new content, and storage use. Check what the person actually receives. In RAG, that includes whether the generated answer is correct and supported by the retrieved evidence. Good retrieval scores alone cannot tell you that.

4. Read the failures

An overall average can hide an important failure. Look at results by language, query type, document format and difficulty. Inspect misses involving negation, quantities, permissions, product versions and exact identifiers.

If the answer was left out when you prepared the documents, a different model cannot retrieve it. If the right passage is found but appears too far down the list, test reranking. If exact identifiers get lost among broadly related results, compare keyword-based or hybrid search.

Each extra component should address a problem you have observed.

The model choice follows you into production

Before choosing a model, I would also consider four practical details.

Long inputs still need careful chunking. A model may accept an entire handbook, but one vector for that handbook may not help you retrieve a small exception inside it. Give each passage enough context to make sense, keep its headings and source references, and test different chunk sizes. Check whether long inputs are being truncated.

Vector size affects storage and cost. More dimensions do not automatically mean better results. With 32-bit floating-point numbers, one million 1,024-dimensional vectors need about 4.1 GB: 1,000,000 × 1,024 × 4 bytes. Metadata, the index and extra copies need more space. Where supported, fewer dimensions or quantisation, which represents values more compactly, can reduce storage. Test the quality of the smaller representation too.

Changing models may mean rebuilding the index. Equal vector lengths do not prove that two models' vectors are comparable. Use a query and document configuration documented to work together, including any supported cross-model compatibility. Google, for example, says Gemini Embedding 2 and gemini-embedding-001 use incompatible embedding spaces, requiring re-embedding when you migrate. Migration documentation

Track the model version, input format, dimensions, normalisation, chunking and index settings together. Keep a way to restore the previous configuration. Avoid gradually mixing incompatible vectors in one collection.

Running costs go beyond the API bill. Include re-embedding, index storage, reranking, hardware and maintenance. Check how long data is kept, where it is stored, and whether the licence allows your intended use. Self-hosting gives you more control and responsibility for operations.

Consider data access too. Embeddings do not anonymise text: researchers have recovered source information under tested attack conditions. Protect the index and check document permissions before content reaches the user or language model. Embedding-inversion research

The decision to make first

For a documentation assistant, begin with models trained for retrieval and passages you have labelled as supporting an answer. For duplicate detection, compare similar kinds of input and consider the cost of an incorrect match. For code or multimodal content, include models trained for those inputs. Keep keyword-based search in the comparison when exact lookups matter.

Then compare a few configurations on your own examples, look at where they fail, and choose the simplest one that meets your quality and operating requirements.

The first question is what counts as a useful match. Once you can describe that, you can test whether a model finds it.

That makes choosing an embedding model an engineering decision you can check, rather than a guess based on a leaderboard.