Cohere released Embed 5 on Wednesday, introducing a system that lets organizations build their search index with the higher-quality Embed 5 Pro model and then run queries against that same index using the faster, less expensive Embed 5 Fast model without creating a duplicate index. The company suggests deploying Pro for indexing operations and Fast for queries, especially in retrieval-augmented generation and agent workflows where delays multiply as systems search the same information repeatedly. There's a quality trade-off, though Cohere's internal testing indicates it's modest.

Across 40 datasets spanning text, images, fused documents, and parsed documents, Fast queries running against a Pro index achieved a score of 98.4 relative to a Pro-to-Pro baseline of 100. Running Fast for both indexing and queries brought that score down to 96.6. Cohere reports that none of the individual datasets displayed a significant decline when Pro and Fast were paired. Pro carries a price of $0.12 per million tokens, while Fast costs $0.08 and provides an average of 2.4 times the document throughput in the company's tests. Both models support six vector dimensions ranging from 256 to 2,048, with float32, int8, and binary formats available. Storage demands shift dramatically at scale: a 2,048-dimensional float32 vector occupies 8 KB, translating to roughly 819 GB for 100 million chunks, while a 1,024-dimensional int8 vector shrinks that to approximately 102 GB, and a 256-dimensional binary vector compresses the same corpus to about 3.2 GB.

The two models share an embedding space, which means teams can switch between them without re-embedding their entire corpus, according to Cohere. Both generate compatible vectors at identical dimensions, and the company states teams can still combine the two when applying Matryoshka truncation or int8 quantization. For retrieval-augmented generation systems that ingest documents less frequently than they search them, Pro can process documents as they enter the index while Fast manages the substantially heavier query traffic. Cohere recommends 1,024-dimensional int8 for most deployments, a configuration that cuts memory and storage requirements while maintaining near full-precision retrieval quality.

The more significant shift in Embed 5 is the capacity to treat indexing and serving as distinct infrastructure decisions, the report explains. The same corpus can be indexed for retrieval quality while the query path gets optimized for throughput and latency, without maintaining two separate data representations. Cohere's 98.4 score suggests the Pro-to-Fast configuration sacrifices relatively little retrieval quality in its tests, though that figure averages across the company's own evaluation suite. Production retrieval-augmented generation and agent systems will still need to benchmark Pro-to-Fast against Pro-to-Pro on their own corpus and query distribution, particularly when retrieval errors can propagate through multiple steps of an agent workflow. Embed 5 Pro and Fast are available through Cohere's API and Model Vault, Microsoft Foundry, and Amazon SageMaker, with private VPC and on-premises deployment supported through vLLM. The split-model architecture addresses a persistent tension in enterprise search deployments, where infrastructure teams have historically paid for retrieval quality even when raw speed would suffice. Organizations building agent systems that loop through the same knowledge base dozens of times per task now face a clearer choice between accuracy and economics at each stage of the workflow.