Cohere Embed 5 Pro Explained: Architecture, Benchmarks, Pricing for RAG
Most retrieval-augmented generation (RAG) systems pay for their embedding model twice: once when a large corpus is indexed, and again, every second, when users ask questions. Those two jobs want opposite things. Indexing wants the best vector money can buy, computed once. Querying wants the cheapest, fastest vector possible, computed millions of times. Cohere Embed 5 Pro is the first widely available embedding release built around that asymmetry, and it shipped on 30 September 2026 alongside a faster sibling, Embed 5 Fast, that lives in the same vector space.
That design matters now because agentic workloads multiply query volume while corpora keep growing, and re-embedding a corpus is the most dreaded migration in the RAG stack. This post separates what Cohere has confirmed from what it has not, works through the storage and cost arithmetic, and lays out a migration plan you can run without gambling production recall.
What this covers: the lineage and design of the Embed 5 family, confirmed specifications, how Matryoshka dimensions and int8 or binary quantization change your storage bill, the company-reported benchmarks and their caveats, deployment and pricing, failure modes, a comparison matrix, and a step-by-step re-embedding migration.
Context and Background
An embedding model maps a piece of content to a point in a high-dimensional space so that semantically similar items land near each other. In a RAG system, documents are split into chunks, each chunk is embedded and stored in a vector index, and at query time the user question is embedded with the same model and matched against the index by nearest-neighbour search. The retrieved chunks are then handed to a large language model (LLM) as grounding context. The embedding model therefore sets a ceiling on everything downstream: if the right passage is not in the top-k candidates, no reranker or LLM can recover it.
Until recently the default assumption was symmetry. You picked one model, used it for both documents and queries, and accepted its latency and price on both sides. If you later wanted a better model, you paid for a full re-index, because vectors from different models are not comparable. That lock-in is why teams stay on embedding models long after better ones appear, and why our own embedding models benchmark across OpenAI, Cohere, Voyage and BGE treats switching cost as a first-class criterion rather than a footnote.
Two other trends shaped the Embed 5 launch. First, enterprise content is increasingly visual: scanned contracts, slide decks, charts, and PDFs where layout carries meaning. Text-only embedding of OCR output loses that structure, which is why visual-document retrieval benchmarks such as ViDoRe have become a competitive battleground. Second, the shift toward agentic RAG, where retrieval agents issue many queries per task, makes query-side latency and cost a bigger share of the total than it was in the single-shot chatbot era.
Cohere’s earlier Embed v3 and v4 generations were strong multilingual models, and Embed 5 continues that line. The announcement positions Pro for frontier quality and Fast for low-latency serving, with the explicit promise that you can mix them. For background on the vendor’s own framing, see Cohere’s Embed 5 announcement and the API changelog entry, which are the primary sources for the specifications below.
What Is Confirmed About Cohere Embed 5 Pro
The short answer: Cohere Embed 5 Pro is a multimodal, multilingual embedding model with a 128K-token context window, six selectable output sizes from 256 to 2,048 dimensions, and float, int8, and binary output formats. It shares an embedding space with Embed 5 Fast, so a corpus indexed with one can be queried with the other. Cohere does not disclose the architecture or training data.
Confirmed specifications
The table below lists only items stated in Cohere’s launch blog or API changelog at the time of writing. Everything else is marked as undisclosed.
| Attribute | Embed 5 Pro | Embed 5 Fast |
|---|---|---|
| API model ID | embed-v5.0-pro |
embed-v5.0-fast |
| Context window | 128K tokens | 128K tokens |
| Output dimensions | 256, 512, 768, 1024, 1536, 2048 | same |
| Output formats | float, int8, binary | same |
| Modalities | text, images, fused text and image | same |
| Languages | 100+ | 100+ |
| Matryoshka support | Yes | Yes |
| Text price | $0.12 per 1M tokens | $0.08 per 1M tokens |
| Image price | $0.40 per 1M tokens | $0.40 per 1M tokens |
| Shared embedding space | Yes | Yes |
What is not disclosed
Cohere has not published a parameter count, the underlying architecture family, the tokenizer or vocabulary, the training corpus, the contrastive training recipe, or the exact Matryoshka training method. The launch material says only that both models support Matryoshka representation learning. Any article that quotes a parameter count or says “built on” a particular backbone is guessing, and you should treat it that way. We label these items undisclosed throughout rather than infer them.
Availability
Embed 5 is available through the Cohere Embed API, through Microsoft Foundry and Amazon SageMaker for both variants, and through Cohere Model Vault, a single-tenant deployment option. Cohere also states that private deployments can run through vLLM. The Foundry catalog lists the Pro model as Cohere-Embed-V5-Pro. We could not confirm regional availability or per-region pricing on the hyperscaler marketplaces, so check your own console before budgeting.
Reference Architecture: Index with Pro, Query with Fast
The architectural idea in the Embed 5 family is not a new transformer trick. It is a contract: two models of different cost and speed are trained so their outputs are interchangeable at a small quality penalty. That contract lets you place the expensive model where it is amortised and the cheap model where it is multiplied.

Figure 1: Reference RAG architecture using Cohere Embed 5 Pro at index time and Embed 5 Fast at query time, joined by a shared embedding space.
The diagram shows two independent paths converging on one vector store. On the left, chunks of text, images, and parsed PDFs flow through Pro and land in the store as int8 or binary vectors. On the right, each user query passes through Fast and is matched against the same store. A reranker then refines the shortlist before the LLM sees it. Because the two models share a space, nothing in the store records which model produced a given vector, and nothing needs to.
Why asymmetric serving pays off
Consider the workload shape. A corpus of 100 million chunks is embedded once, then incrementally as documents change. The query stream is open-ended and latency-sensitive, and in agentic systems a single user task may trigger dozens of retrieval calls. Cohere reports that Fast delivers an average of 2.4 times the throughput of Pro across context sizes, measured at 377.3 versus 159.7 documents per second. That is a company-reported measurement and the hardware and batch settings are not part of the headline claim, so treat it as directional.
The economics follow the same logic. At list prices, text costs $0.12 per million tokens on Pro and $0.08 on Fast, a one-third discount. The discount on Fast looks modest, but throughput is where the real saving appears if you self-host or pay for dedicated capacity: 2.4 times the documents per second per accelerator means roughly 58 percent fewer accelerators for the same query load, assuming linear scaling, which is a simplification you should verify in your own environment.
The compatibility contract and its price
Shared-space compatibility is not free. Cohere reports that mismatched pairings lose a small amount of quality on a normalised nDCG@10 average across 40 development datasets: 1.6 percent when Fast queries run against a Pro-indexed corpus, and 2.7 percent when Pro queries run against a Fast-indexed corpus. The baseline is Pro queries on a Pro corpus. Read that carefully. The cheap direction, Fast queries over a Pro index, is also the smaller loss, which is exactly the configuration the architecture recommends.
Two cautions apply. These are averages over Cohere’s own development datasets, so your domain may sit well above or below them. And a 1.6 percent average can hide a long tail: a handful of query types may lose far more while most lose nothing. That is why the migration plan later in this post insists on running your own evaluation set before cutting over.
Mixed-modality indexing
Embed 5 accepts text, images, and fused text-plus-image inputs in the same space. Practically, that means a slide or a scanned page can be embedded as an image, as extracted text, or as both fused, and queried with plain-language text. Images are priced at $0.40 per million tokens on both models, more than three times the Pro text rate, so the choice between embedding raw page images and embedding parsed text is partly a cost decision. Cohere has not said in the launch material how images are tokenised, so we cannot give you a tokens-per-page figure. Measure it on a sample before estimating a bill.
Deeper Analysis: Dimensions, Quantization, Benchmarks, and Cost
This section turns the specification into numbers: how big your index will be, how the Matryoshka and quantization options interact, what the published benchmark scores do and do not tell you, and what a realistic bill looks like.
Matryoshka dimensions and quantization
Matryoshka representation learning trains a vector so that its leading dimensions carry the most information. You can truncate a 2,048-dimension vector to its first 1,024, 512, or 256 values and still get a usable embedding, with graceful rather than catastrophic quality loss. Embed 5 exposes six sizes directly through the API, so you request the size you want instead of truncating by hand. Our deeper explainer on Matryoshka embeddings and adaptive retrieval covers the training objective; here we focus on what it does to storage.
Quantization is the second lever. Float vectors use 32 bits per dimension, int8 uses 8, and binary uses 1. Cohere’s own example makes the spread concrete: a 2,048-dimension float32 vector takes 8 KB, a 1,024-dimension int8 vector takes 1 KB, and a 256-dimension binary vector takes 32 bytes, a 256-fold reduction. For 100 million chunks that moves raw vector storage from roughly 819 GB to 3.2 GB. Cohere states that int8 retains near-full-precision retrieval quality, but it does not publish a per-dimension, per-format quality table in the launch material, so the quality of the aggressive corners (256-dimension binary) is something you must measure.

Figure 2: How Matryoshka truncation and quantization combine, and how a two-stage search uses a tiny coarse vector then rescoring on a larger one.
The diagram shows the pattern most large deployments converge on. A compact binary vector drives a fast coarse scan that produces a shortlist, and a larger int8 vector rescores only that shortlist. The compact vectors keep the hot index in memory; the larger ones can live on cheaper storage and are fetched only for candidates. Whether your vector database supports storing two representations per record is the practical question, and it is worth checking before you design around it.
Storage arithmetic you can reproduce
The following table uses the byte arithmetic implied by the formats above. It counts raw vector bytes only. Index structures such as HNSW graph links, metadata, replicas, and filtering columns add overhead that varies by database, so treat these as lower bounds.
| Configuration | Bytes per vector | 100M vectors raw | Relative to 2048 float |
|---|---|---|---|
| 2048 dims, float32 | 8,192 | about 819 GB | 1x |
| 1536 dims, float32 | 6,144 | about 614 GB | 0.75x |
| 1024 dims, float32 | 4,096 | about 410 GB | 0.5x |
| 1024 dims, int8 | 1,024 | about 102 GB | 0.125x |
| 512 dims, int8 | 512 | about 51 GB | 0.0625x |
| 256 dims, int8 | 256 | about 26 GB | 0.031x |
| 256 dims, binary | 32 | about 3.2 GB | 0.0039x |
Two things stand out. First, quantization saves more per step than truncation: going from float32 to int8 is a fixed 4x, while halving dimensions is 2x. Second, the interesting zone for most teams is 1024-dimension int8 at 1 KB per vector. It cuts memory by 8x against the full-size float vector and, per Cohere’s statement, keeps near-full-precision quality. That is the configuration we would test first. These numbers are our arithmetic from Cohere’s published formats, not measured recall.
What the benchmarks say
Cohere reports results using a metric it calls RCP-nDCG@10, described as evaluating retrieved documents against query-specific relevance criteria rather than only a limited set of fixed labels. The aim is to avoid penalising a model for returning a genuinely relevant document that the original annotators never labelled. It is a reasonable idea, and it also means the headline numbers are not directly comparable to published nDCG@10 scores from other leaderboards.
| Benchmark | Embed 5 Pro | Embed 5 Fast | Notes |
|---|---|---|---|
| ViDoRe V3 | 85.8 | 84.5 | visually rich enterprise documents |
| FinanceBench | 80.1 | 80.0 | financial document QA retrieval |
| FinQA | 90.0 | 88.8 | financial numerical reasoning retrieval |
| Parsed PDFs average | 84.8 | 83.4 | suite average |
On ViDoRe V3, Cohere’s comparison lists Voyage 4 Large at 83.7, Google Gemini Embedding 2 at 83.2, and OpenAI text-embedding-3-large at 75.5. All of these figures originate from Cohere’s launch material and secondary coverage of it, and we did not find an independent reproduction at the time of writing. We have not verified them against the public ViDoRe leaderboard.
Read the table for its shape rather than its decimals. Pro leads Fast by 1.3 points on ViDoRe V3, 0.1 on FinanceBench, 1.2 on FinQA, and 1.4 on parsed PDFs. The gap is small enough that for many corpora Fast alone may be sufficient, and large enough that Pro is justified where you index once and recall errors are expensive. The 2.1-point lead over the nearest competitor on ViDoRe V3 is within the range where prompt, chunking, and reranker choices can swing a real system by more.
Cost model with worked numbers
The figures below are illustrative inputs applied to Cohere’s list prices; your chunk sizes and query lengths will differ.
Suppose you index 100 million chunks averaging 300 tokens, which is 30 billion tokens. On Pro text at $0.12 per million that is $3,600 for the initial embedding. On Fast at $0.08 it would be $2,400. The one-time difference is $1,200, which is small against engineering and evaluation time.
Now the query side. Suppose an agentic product issues 10 million retrieval queries a month at 40 tokens each, 400 million tokens. On Pro that is $48 a month; on Fast, $32. At this scale the API embedding cost of queries is trivial. The real savings from Fast show up in latency and in self-hosted capacity, not in the API invoice. If your volume is a thousand times higher, the monthly difference becomes meaningful, but even then storage and the vector database dominate. A 100-million-vector index at 1 KB per vector is about 100 GB before overhead, and the memory to serve it costs far more than the embedding calls.
The practical conclusion: do not choose between Pro and Fast on API price. Choose on latency budget and quality margin, and let the shared space let you change your mind later.

Figure 3: Sequence for migrating a production RAG index to Embed 5 Pro with a shadow evaluation and an alias cutover.
The sequence above is the migration pattern we recommend, and we walk through it step by step in the recommendations section. The critical property is that the old index keeps serving traffic until the evaluation harness proves the new one is at least as good on your own queries.
Access, Deployment, and Calling the API
Embed 5 reaches production through four routes, and the choice mostly follows your data-governance constraints rather than model quality, since the weights behind each route are the same family.
Routes and when to use each
The Cohere Embed API is the fastest way to start and the right default for non-sensitive corpora. Microsoft Foundry and Amazon SageMaker let you consume the models inside an existing cloud contract and network boundary, which often unlocks procurement and data-residency approvals that a direct vendor API would not. Cohere Model Vault provides single-tenant deployment, and vLLM support is stated for private deployments. We found no published license terms for self-hosted weights, no stated GPU memory requirement, and no published latency percentiles, so any capacity plan for a private deployment needs a proof-of-concept before commitment. Those are gaps, not hidden facts we are withholding.
A minimal indexing and query example
The snippet below follows Cohere’s established embed API pattern with the new model IDs. Parameter names for dimension and format selection follow the Embed v4-era conventions; confirm them against the current API reference before use, because we verified the model IDs and supported values from the changelog but not every parameter name.
import cohere
co = cohere.ClientV2("YOUR_API_KEY")
# Index time: use the higher-quality model, compact int8 vectors
doc_resp = co.embed(
model="embed-v5.0-pro",
input_type="search_document",
texts=["Chunk of a maintenance manual...", "Another chunk..."],
embedding_types=["int8"],
output_dimension=1024,
)
doc_vectors = doc_resp.embeddings.int8
# Query time: use the faster model in the same embedding space
q_resp = co.embed(
model="embed-v5.0-fast",
input_type="search_query",
texts=["torque spec for the main bearing cap"],
embedding_types=["int8"],
output_dimension=1024,
)
query_vector = q_resp.embeddings.int8[0]
Three details matter in practice. Keep output_dimension and the numeric format identical on both sides, because the shared space does not remove the need for matching shapes. Use the document and query input types consistently, since asymmetric embedding models are trained to expect them. And store the model ID and dimension alongside every vector as metadata, so that a future migration can find and rebuild exactly the records that need it.
Choosing a path by workload

Figure 4: Choosing Embed 5 Pro, Embed 5 Fast, or a private deployment based on workload shape.
The decision tree reduces to three questions. Is the corpus large and mostly static, with errors expensive? Index with Pro. Are queries interactive or agent-driven, with tight latency budgets? Query with Fast. Is the data restricted? Use Model Vault or a vLLM private deployment, and let the shared space keep the other two choices open. Because all branches converge on one embedding space, the decisions are independent, and you can revisit each without re-indexing, within the quality cost stated above.
How Embed 5 Pro Compares
Embedding choice is rarely a pure leaderboard decision. The matrix below compares Embed 5 Pro with peers named in Cohere’s own comparison, using only attributes we could source. Where a peer attribute is unknown to us, the cell says so rather than guessing.
| Use case | Embed 5 Pro | Voyage 4 Large | Gemini Embedding 2 | OpenAI text-embedding-3-large |
|---|---|---|---|---|
| Scanned and visual enterprise documents | Strong reported score of 85.8 on ViDoRe V3 | 83.7 reported | 83.2 reported | 75.5 reported |
| Cost-tuned storage at scale | Six dimensions plus int8 and binary | Not verified here | Not verified here | Not verified here |
| Query-time latency tuning | Fast sibling in the same space | Not verified here | Not verified here | Not verified here |
| Private or single-tenant deployment | Model Vault, vLLM, Foundry, SageMaker | Not verified here | Not verified here | Not verified here |
The honest reading is narrow. On the one benchmark where the competing numbers are published side by side, Pro leads, all of it self-reported. On deployment flexibility, the Pro and Fast pairing is the differentiated feature, because it gives you a supported way to decouple index quality from query latency. For a wider survey including open-weight options, see our open-source embedding models benchmark and the earlier cross-vendor comparison linked above.
Limitations, Failure Modes, and What Goes Wrong
Every embedding upgrade arrives with a launch-day halo. These are the failure modes we would plan for.
Benchmark transfer. ViDoRe V3, FinanceBench, and FinQA measure retrieval on specific document types. If your corpus is code, chat logs, or short product titles, the headline gains may not appear. Build a golden set of at least a few hundred real queries with judged relevant passages, and weight it by traffic. A model that wins public benchmarks can lose on a domain-specific acronym-heavy corpus where lexical matching matters.
The long tail inside the average. The reported 1.6 and 2.7 percent mismatch losses are averages across 40 development datasets. A mixed deployment can degrade unevenly: short keyword-style queries may suffer more than natural-language questions, because a faster model has less capacity to disambiguate them. Segment your evaluation by query length and intent before trusting the mean.
The long context trap. A 128K-token window is a capability, not a recommendation. Embedding a whole long document into one vector averages away detail, so a query about one clause retrieves a vector dominated by the other ninety. Chunking remains necessary for precision, and the long window is best used for chunks that need surrounding context, such as tables with headers or clauses with definitions. Cohere has not published a quality-versus-length curve for Embed 5, so test it.
Aggressive compression. Binary 256-dimension vectors are a 256x reduction and they are not free. Use them as a coarse first stage with int8 or float rescoring, as in Figure 2, rather than as the only representation. Recall at the corner of the compression range is something Cohere has not published, so measure it with your own data.
Mixed-model drift. The shared space makes mixing safe at query time, but it invites quiet inconsistency: one team embeds with Pro and another with Fast into the same collection, without recording which. The aggregate quality then sits somewhere between the two figures and nobody can explain a regression. Tag every vector with model ID, dimension, format, and embedding date.
Vendor and metric lock-in. The shared space is shared between Cohere’s own models, not with other vendors. Adopting Embed 5 reduces re-embedding cost between Pro and Fast, but it does not give you a path to Voyage or an open-weight model without a full re-index. Keep the original text for every chunk so that a future migration is a recompute, not a re-ingestion.
Hybrid retrieval still matters. Dense embeddings underperform on exact identifiers, part numbers, and rare tokens. Pair Embed 5 with a lexical retriever and a reranker, as we describe in our production hybrid search ADR. Better dense recall narrows the gap but does not erase it.
Graph and structure. If your questions are multi-hop, such as which supplier of a component also supplied a failed assembly, a stronger vector model alone will not help. That is the territory of GraphRAG retrieval patterns, where embeddings seed a graph traversal rather than answer the question directly.
Practical Recommendations: A Migration Plan
If you are on an older embedding model today, the question is not whether Embed 5 Pro is better in the abstract. It is whether the quality gain on your queries repays a re-embedding project. Here is a sequence that keeps the answer evidence-based and the rollback trivial.
First, freeze a baseline. Build an evaluation set from real query logs, with relevance judgments from domain experts or a calibrated LLM judge, and record recall at k, mean reciprocal rank, and nDCG@10 on your current index. Without a baseline, every later claim is anecdote.
Second, embed a sample, not the whole corpus. Take a representative one to five percent slice, embed it with Pro at two or three dimension and format combinations, and compare against the baseline on the same slice. The sample cost is a rounding error at $0.12 per million tokens, and it tells you whether the full project is worthwhile.
Third, backfill into a new index alongside the old one. Use a collection or alias name that includes the model version. Keep the old index serving. Run the backfill at off-peak rates, batching to respect API limits, and store model ID, dimension, and format with each record.
Fourth, shadow the query path. Send a copy of live queries to the new index using Fast, and compare top-k overlap and downstream answer quality offline. This is where you measure the 1.6 percent mismatch figure on your own data instead of trusting Cohere’s average.
Fifth, cut over with an alias, then soak. Flip the alias, watch retrieval and answer-quality dashboards, and keep the old index for a defined window before retirement. Because rollback is an alias flip, the decision is cheap to reverse.
A short checklist to close:
- Baseline recall, MRR, and nDCG@10 recorded on a golden query set.
- Sample embedded at 1024-dimension int8 and one alternative, results compared.
- Model ID, dimension, and format stored as metadata on every vector.
- Pro for indexing, Fast for queries, validated by shadow traffic.
- Binary vectors used only as a first stage, with rescoring.
- Hybrid lexical retrieval and a reranker retained.
- Original chunk text archived so the next migration is a recompute.
- Old index retained until the soak window passes.
Finally, keep the economics in proportion. The one-time embedding cost for a 100-million-chunk corpus in our worked example is a few thousand dollars at list price. The larger costs are the engineering hours, the evaluation set, and the vector database capacity, so spend your attention there.
Frequently Asked Questions
What is Cohere Embed 5 Pro?
Cohere Embed 5 Pro is the higher-quality model in Cohere’s Embed 5 family, released on 30 September 2026. It embeds text, images, and fused text-and-image inputs across more than 100 languages, accepts up to 128K tokens, and offers 256 to 2,048 output dimensions in float, int8, or binary formats. Its API ID is embed-v5.0-pro. Cohere has not disclosed its architecture or parameter count.
What is the difference between Embed 5 Pro and Embed 5 Fast?
Both share specifications and an embedding space. Pro is optimised for quality and is priced at $0.12 per million text tokens; Fast is optimised for speed, priced at $0.08, and Cohere reports about 2.4 times the throughput. Reported benchmark gaps are small, for example 85.8 versus 84.5 on ViDoRe V3. Cohere recommends Pro for indexing and Fast for interactive queries and agent workflows.
Can I index with Pro and query with Fast?
Yes. Cohere states that Pro and Fast share an embedding space, so a corpus indexed with one can be queried with the other. The company reports average losses of 1.6 percent for Fast queries against a Pro corpus and 2.7 percent for Pro queries against a Fast corpus, relative to Pro on Pro, across 40 development datasets. Validate this on your own queries before relying on it.
How much does Cohere Embed 5 Pro cost?
At the time of writing, Cohere lists Embed 5 Pro at $0.12 per million text tokens and $0.40 per million image tokens. Embed 5 Fast is $0.08 for text and $0.40 for images. Embedding 30 billion tokens with Pro would cost about $3,600 at list price. Marketplace pricing on Microsoft Foundry and Amazon SageMaker, and private deployment costs, may differ, and we could not verify them.
Do I need to re-embed my data to adopt Embed 5?
If you are moving from a different model, including Cohere’s earlier versions, yes. Vectors from different models are not comparable, so the whole corpus must be re-embedded into a new index. The shared space only removes that requirement between Embed 5 Pro and Embed 5 Fast. Plan a parallel index, evaluate on your own queries, and cut over with an alias.
Is it worth using binary or int8 embeddings?
Often, yes, with care. Cohere says int8 retains near-full-precision retrieval quality, and a 1,024-dimension int8 vector uses 1 KB versus 8 KB for a 2,048-dimension float32 vector. Binary 256-dimension vectors shrink storage 256 times but are best used as a coarse first stage with rescoring. Cohere has not published a full quality table by format, so test on your data.
Further Reading
- Embedding models benchmark: OpenAI, Cohere, Voyage and BGE for the cross-vendor baseline.
- Matryoshka embeddings and adaptive retrieval architecture for the truncation mechanics behind the dimension options.
- Agentic RAG architecture and retrieval agents for why query volume is rising.
- GraphRAG hybrid retrieval and knowledge graph pattern for multi-hop questions that vectors alone cannot answer.
- GraphRAG knowledge graph retrieval-augmented generation architecture for the foundational pattern.
- Cohere: Introducing Embed 5 and the Embed 5 API changelog, the primary sources for every specification above.
By Riju — about
