GPAI Ledger The public record of EU AI Act training-data summaries

GPAI LedgerModel catalog (developer docs) (Cohere) › Capture 19 Aug 2026

Model catalog (developer docs) — capture 20260819T150449Z

ProviderCohere
Targetwatched page — https://docs.cohere.com/docs/models
Fetched (UTC)2026-08-19T15:04:49Z
Stored fileb5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea.html.txt (812,630 bytes) (served with a .txt suffix so the captured page cannot run scripts on this site; bytes are identical — the SHA-256 verifies against this file)
SHA-256b5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea
OpenTimestamps proofb5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea.html.ots (calendar-attested; anchored in bitcoin over time)
WaybackWayback snapshot, 2026-08-19 15:04 UTC
Prior capture of this target— first capture of this target

Verify: sha256sum b5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea.html.txt must equal the hash above (the filename IS the expected hash); ots verify b5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea.html.ots -f b5c8ede22b7c63a54209f4cde30f7bef29ea58f88bb31617998c9353dfd8aeea.html.txt (opentimestamps.org) proves the capture time (fresh proofs report 'pending' until bitcoin-anchored, typically within a day).

Extracted text

Machine-extracted text (layout may be lost; the authoritative content is the stored file above).

An Overview of Cohere's Models | Cohere
For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
docs
v2 API
v2 API
Search
/
Ask AI
DASHBOARD
PLAYGROUND
DOCS
COMMUNITY
LOG IN
Guides and concepts
API Reference
Release Notes
LLMU
Cookbooks
Search
/
Ask AI
v2 API
v2 API
Guides and concepts
API Reference
Release Notes
LLMU
Cookbooks
Get Started
Welcome to Cohere
Cohere Platform
Overview
Installation
Creating a client
Quickstart
Playground
FAQs
Model Vault
Overview
Quickstart
Deploy & manage
Operate & observe
Standard Vault
Encrypted Vault
Model Vault with North
Models
An Overview of Cohere's Models
Audio
Aya
Command
Embed
North
Rerank
Text Generation
Introduction to Text Generation at Cohere
Using the Chat API
Reasoning
Image Inputs
Streaming Responses
Structured Outputs
Predictable Outputs
Advanced Generation Parameters
Retrieval Augmented Generation (RAG)
Tool Use
Tokens and Tokenizers
Summarizing Text
Safety Modes
Embeddings (Vectors, Search, Retrieval)
Introduction to Embeddings at Cohere
Semantic Search with Embeddings
Multimodal Embeddings
Batch Embedding Jobs
Reranking
Going to Production
API Keys and Rate Limits
Going Live
Deprecations
How Does Cohere's Pricing Work?
Integrations
Integrating Embedding Models with Other Tools
Cohere and LangChain
LlamaIndex and Cohere
Deployment Options
Overview
SDK Compatibility
Private Deployment
Cloud AI Services
Tutorials
Cookbooks
LLM University
Build Things with Cohere!
Agentic RAG
Cohere on Azure
Responsible Use
Cohere Trust Center
Security
Usage Policy
Command A Technical Report
Command R and Command R+ Model Card
Cohere Web Crawlers
Cohere Labs
Cohere Labs Acceptable Use Policy
More Resources
Cohere Toolkit
Datasets
Improve Cohere Docs
Welcome to Cohere
Overview
Installation
Creating a client
RAG
Reranking
Semantic Search
Text Generation
Tool Use & Agents
Transcribing Audio
Playground
FAQs
Overview
Quickstart
Home Page
Creating a Vault
Managing Vaults
Monitoring
Overview
Supported Models
Calling the API
Pricing
Overview
Supported Models
Calling the API
Confidential Computing Primer
Security Model
Remote Attestation
Verifying Your Deployment
Encryption & Key Management
Compliance
Pricing
FAQ
Model Vault with North
An Overview of Cohere's Models
Cohere Transcribe
Cohere Transcribe Arabic
Aya
Aya Vision
Aya Expanse
Tiny Aya
Command A+
Command A
Command A Reasoning
Command A Translate
Command A Vision
Command R7B
Command R+
Command R
Embed
North Mini Code
Rerank
Introduction to Text Generation at Cohere
Using the Chat API
Reasoning
Image Inputs
Streaming Responses
Structured Outputs
Parameter Types in Structured Outputs (JSON)
Predictable Outputs
Advanced Generation Parameters
Basic usage
End-to-end example
Streaming
Citations
Tool Use
Basic usage
Usage patterns
Parameter types
Streaming
Citations
Tokens and Tokenizers
Summarizing Text
Safety Modes
Introduction to Embeddings at Cohere
Semantic Search with Embeddings
Multimodal Embeddings
Batch Embedding Jobs
Rerank Overview
Rerank Best Practices
API Keys and Rate Limits
Going Live
Deprecations
How Does Cohere's Pricing Work?
Integrating Embedding Models with Other Tools
Elasticsearch and Cohere
MongoDB and Cohere
Redis and Cohere
Haystack and Cohere
Pinecone and Cohere
Weaviate and Cohere
Open Search and Cohere
Vespa and Cohere
Qdrant and Cohere
Milvus and Cohere
Zilliz and Cohere
Chroma and Cohere
Cohere and LangChain
Chat on LangChain
Embed on LangChain
Rerank on LangChain
Tools on LangChain
LlamaIndex and Cohere
Overview
SDK Compatibility
Overview
Setting Up
Model Deployment
Model Deployment - AWS
Usage
Cohere on AWS
Amazon Bedrock
Amazon SageMaker
Deploy Your Own Finetuned Command-R-0824 Model from AWS Marketplace
Cohere on Azure
Cohere on Oracle Cloud Infrastructure (OCI)
Cookbooks
LLM University
Build Things with Cohere!
Cohere Text Generation Tutorial
Building a Chatbot with Cohere
Semantic Search with Cohere
Reranking with Cohere
RAG with Cohere
Building an Agent with Cohere
Agentic RAG
Routing Queries to Data Sources
Generating Parallel Queries
Performing Tasks Sequentially
Generating Multi-Faceted Queries
Querying Structured Data (Tables)
Querying Structured Data (SQL)
Cohere on Azure
Text Generation
Semantic Search
Reranking
Retrieval Augmented Generation (RAG)
Tool Use & Agents
Usage Policy
Command R and Command R+ Model Card
Cohere Web Crawlers
Cohere Labs Acceptable Use Policy
Cohere Toolkit
Datasets
Improve Cohere Docs
DASHBOARD
PLAYGROUND
DOCS
COMMUNITY
LOG IN
Light
On this page
What can These Models Be Used For?
Command
Using Command Models on Different Platforms
Embed
Using Embed Models on Different Platforms
Rerank
Using Rerank Models on Different Platforms
Audio
Using Audio Models on Different Platforms
Aya
Scroll to top
Models
An Overview of Cohere's Models
Copy page
Cohere has a variety of models that cover many different use cases. If you need more customization, you
can
tune your prompts
 to adjust its behavior to your specific
use case.

Cohere models are currently available on the following platforms:

Cohere’s proprietary platform

Amazon SageMaker

Amazon Bedrock

Microsoft Azure

Oracle GenAI Service

At the end of each major section below, you’ll find technical details about how to call a given model on a
particular platform.

What can These Models Be Used For?

In this section, we’ll provide some high-level context on Cohere’s offerings, and what the strengths of each
are.

The Command family of models includes
Command A+
,
Command A
,

Command R7B
,
Command A Translate
,

Command A Reasoning
,
Command A Vision
,

Command R+
,
Command R
, and
Command
.
Together, they are the text-generation LLMs powering tool-using agents,

retrieval augmented generation
 (RAG), translation, copywriting, and
similar use cases. They work through the
Chat
 endpoint, which can be used with or without
RAG.

Rerank
 is the fastest way to inject the intelligence of a language model into an existing search system. It can be accessed via the
Rerank
 endpoint.

Embed
 improves the accuracy of search, classification, clustering, and RAG results. It powers the
Embed
 endpoint.

Cohere Transcribe
 is Cohere’s dedicated audio transcription model for automatic speech
recognition (ASR). It powers the
Audio Transcriptions
 endpoint.

Cohere Transcribe Arabic
 is a version of the model optimized for Arabic-language
audio.

The
Aya
 family of models are aimed at expanding the number of languages covered by
generative AI. Aya Expanse covers 23 languages, and Aya Vision is fully multimodal, allowing you to pass
in images and text and get a single coherent response. Both are available on the
Chat

endpoint.

Command

Command is Cohere’s default generation model that takes a user instruction (or command) and generates text
following the instruction. Our Command models also have conversational capabilities, meaning they are
well-suited for chat applications, and Command A Vision can interact with
image inputs
.

Model Name
Status
Description
Modality
Context Length
Maximum Output Tokens
Endpoints
command-a-plus-05-2026
Live
Command A+ offers the last model in the Command A family, while being Cohere’s first Mixture of Experts model,  simultaneously combining vision input support, agentic, reasoning, and world-class translation capabilities into single model weights. It also can fit on 1 x B200 or 2 x H100 GPUs, while providing significant latency and throughput improvements over Command A Reasoning, making it an ideal model for any enterprise to deploy at scale.
Text, Images
128k
64k
Chat
command-a-03-2025
Live
Command A is our most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.
Text
256k
8k
Chat
command-r7b-12-2024
Live
command-r7b-12-2024
 is a small, fast update delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning and multiple steps.
Text
128k
4k
Chat
command-a-translate-08-2025
Live
Command A Translate is Cohere’s state of the art machine translation model, excelling at a variety of translation tasks on 23 languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Chinese, Arabic, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, Persian.
Text
8K
8k
Chat
command-a-reasoning-08-2025
Live
Command A Reasoning is Cohere’s first reasoning model, able to ‘think’ before generating an output in a way that allows it to perform well in certain kinds of nuanced problem-solving and agent-based tasks in 23 languages.
Text
256k
32k
Chat
command-a-vision-07-2025
Live
Command A Vision is our first model capable of processing images, excelling in enterprise use cases such as analyzing charts, graphs, and diagrams, table understanding, OCR, document Q&A, and object detection. It officially supports English, Portuguese, Italian, French, German, and Spanish.
Text, Images
128K
8K
Chat
command-r-08-2024
Live
command-r-08-2024
 is an update of the Command R model, delivered in August 2024. Find more information in the
changelog
Text
128k
4k
Chat
command-r-plus-08-2024
Live
command-r-plus-08-2024
 is an update of the Command R+ model, delivered in August 2024. Find more information in the
changelog
Text
128k
4k
Chat
command-r-03-2024
Deprecated Sept 15, 2025
Command R is an instruction-following conversational model that performs language tasks at a higher quality, more reliably, and with a longer context than previous models. It can be used for complex workflows like code generation, retrieval augmented generation (RAG), tool use, and agents.
Text
128k
4k
Chat
command-r-plus-04-2024
Deprecated Sept 15, 2025
Command R+ is an instruction-following conversational model that performs language tasks at a higher quality, more reliably, and with a longer context than previous models. It is best suited for complex RAG workflows and multi-step tool use.
Text
128k
4k
Chat
command-r-plus
Deprecated Sept 15, 2025
Alias for
command-r-plus-04-2024
Text
128k
4k
Chat
command-r
Deprecated Sept 15, 2025
Alias for
command-r-03-2024
Text
128k
4k
Chat
command-light
Deprecated Sept 15, 2025
A smaller, faster version of
command
. Almost as capable, but a lot faster.
Text
4k
4k
Chat
command
Deprecated Sept 15, 2025
An instruction-following conversational model that performs language tasks with high quality, more reliably and with a longer context than our base generative models.
Text
4k
4k
Chat

Using Command Models on Different Platforms

In this table, we provide some important context for using Cohere Command models on Amazon Bedrock, Amazon SageMaker, and more.

Model Name
Amazon Bedrock Model ID
Amazon SageMaker
Azure AI Foundry
Oracle OCI Generative AI Service
command-a-plus-05-2026
N/A
N/A
’coherelabs-command-a-plus-05-2026-w4a4’
N/A
command-a-03-2025
(Coming Soon)
Unique per deployment
Unique per deployment
cohere.command-a-03-2025
command-r7b-12-2024
N/A
N/A
N/A
N/A
command-r-plus
cohere.command-r-plus-v1:0
Unique per deployment
Unique per deployment
cohere.command-r-plus v1.2
command-r
cohere.command-r-v1:0
Unique per deployment
Unique per deployment
cohere.command-r-16k v1.2
command
cohere.command-text-v14
N/A
N/A
cohere.command v15.6
command-nightly
N/A
N/A
N/A
N/A
command-light
cohere.command-light-text-v14
N/A
N/A
cohere.command-light v15.6
command-light-nightly
N/A
N/A
N/A
N/A

Embed

These models can be used to generate embeddings from text or classify it based on various parameters. Embeddings can be used for estimating semantic similarity between two sentences, choosing a sentence which is most likely to follow another sentence, or categorizing user feedback. The Representation model comes with a variety of helper functions, such as for detecting the language of an input.

Model Name
Description
Modalities
Dimensions
Context Length
Similarity Metric
Endpoints
embed-v4.0
A model that allows for text and images to be classified or turned into embeddings
Text, Images, Mixed texts/images (i.e. PDFs)
One of ‘[256, 512, 1024, 1536 (default)]‘
128k
Cosine Similarity, Dot Product Similarity, Euclidean Distance
Embed
,
Embed Jobs
embed-english-v3.0
A model that allows for text to be classified or turned into embeddings. English only.
Text, Images
1024
512
Cosine Similarity
Embed
,
Embed Jobs
embed-english-light-v3.0
A smaller, faster version of
embed-english-v3.0
. Almost as capable, but a lot faster. English only.
Text, Images
384
512
Cosine Similarity
Embed
,
Embed Jobs
embed-multilingual-v3.0
Provides multilingual classification and embedding support.
See supported languages here.
Text, Images
1024
512
Cosine Similarity
Embed
,
Embed Jobs
embed-multilingual-light-v3.0
A smaller, faster version of
embed-multilingual-v3.0
. Almost as capable, but a lot faster. Supports multiple languages.
Text, Images
384
512
Cosine Similarity
Embed
,
Embed Jobs

Using Embed Models on Different Platforms

In this table, we provide some important context for using Cohere Embed models on Amazon Bedrock, Amazon SageMaker, and more.

Model Name
Amazon Bedrock Model ID
Amazon SageMaker
Azure AI Foundry
Oracle OCI Generative AI Service
embed-v4.0
(Coming Soon)
Unique per deployment
cohere-embed-v-4-plan
(Coming Soon)
embed-english-v3.0
cohere.embed-english-v3
Unique per deployment
Unique per deployment
cohere.embed-english-image-v3.0
 (for images),
cohere.embed-english-v3.0
 (for text)
embed-english-light-v3.0
N/A
Unique per deployment
N/A
cohere.embed-english-light-image-v3.0
 (for images),
cohere.embed-english-light-v3.0
 (for text)
embed-multilingual-v3.0
cohere.embed-multilingual-v3
Unique per deployment
Unique per deployment
cohere.embed-multilingual-image-v3.0
 (for images),
cohere.embed-multilingual-v3.0
 (for text)
embed-multilingual-light-v3.0
N/A
Unique per deployment
N/A
cohere.embed-multilingual-light-image-v3.0
 (for images),
cohere.embed-multilingual-light-v3.0
 (for text)
embed-english-v2.0
N/A
Unique per deployment
N/A
N/A
embed-english-light-v2.0
N/A
Unique per deployment
N/A
cohere.embed-english-light-v2.0
embed-multilingual-v2.0
N/A
Unique per deployment
N/A
N/A

Rerank

The Rerank model can improve created models by re-organizing their results based on certain parameters. This can be used to improve search algorithms.

Model Name
Description
Modalities
Context Length
Endpoints
rerank-v4.0-pro
A multilingual model that allows for re-ranking English and non-english documents and semi-structured data (JSON). This model is better suited for state-of-the-art quality and complex use-cases than its
fast
 variant.
Text
32k
Rerank
rerank-v4.0-fast
A light version of
rerank-v4.0-pro
, this is a multilingual model that allows for re-ranking English and non-english documents and semi-structured data (JSON). This model is better suited for low latency and high throughput use-cases than its
pro
 variant.
Text
32k
Rerank
rerank-v3.5
A model that allows for re-ranking English Language documents and semi-structured data (JSON). This model has a context length of 4096 tokens.
Text
4k
Rerank
rerank-english-v3.0
A model that allows for re-ranking English Language documents and semi-structured data (JSON). This model has a context length of 4096 tokens.
Text
4k
Rerank
rerank-multilingual-v3.0
A model for documents and semi-structure data (JSON) that are not in English. Supports the same languages as
embed-multilingual-v3.0
. This model has a context length of 4096 tokens.
Text
4k
Rerank

Using Rerank Models on Different Platforms

In this table, we provide some important context for using Cohere Rerank models on Amazon Bedrock, SageMaker, and more.

Model Name
Amazon Bedrock Model ID
Amazon SageMaker
Azure AI Foundry
Oracle OCI Generative AI Service
rerank-v4.0-pro
N/A
Unique per deployment
cohere-rerank-v4-pro
N/A
rerank-v4.0-fast
N/A
Unique per deployment
cohere-rerank-v4-fast
N/A
rerank-v3.5
cohere.rerank-v3-5:0
Unique per deployment
Cohere-rerank-v3.5
cohere.rerank.3-5
rerank-english-v3.0
N/A
Unique per deployment
Cohere-rerank-v3-english
N/A
rerank-multilingual-v3.0
N/A
Unique per deployment
Cohere-rerank-v3-multilingual
N/A

Rerank accepts full strings rather than tokens, so the token limit works a little differently. Rerank will automatically chunk documents longer than 510 tokens, and there is therefore no explicit limit to how long a document can be when using rerank. See our
best practice guide
 for more info about formatting documents for the Rerank endpoint.

Audio

Cohere Transcribe is our dedicated model for audio-in, text-out automatic speech recognition (ASR) workloads. For Arabic-language transcription, use Cohere Transcribe Arabic for best-in-class performance.

Model Name
Status
Description
Maximum file size
Endpoints
cohere-transcribe-03-2026
Live
Open source research release focused on high-accuracy, multilingual speech transcription.
25MB
Audio Transcriptions

Using Audio Models on Different Platforms

Cohere Transcribe is available on Microsoft Foundry under ‘coherelabs-cohere-transcribe-03-2026’. Cohere Transcribe Arabic is not yet available on other platforms.

Aya

Aya
 is a family of multilingual large language models designed to expand the number of languages covered by generative AI for purposes of research and to better-serve minority linguistic communities.

The 32-billion parameter Aya Expanse offering is optimized to perform well in these 23 languages: Arabic,
Chinese (simplified & traditional), Czech, Dutch, English, French, German, Greek, Hebrew, Hebrew, Hindi, Indonesian,
Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Turkish, Ukrainian, and
Vietnamese.

The 32-billion parameter Aya Vision model is a state-of-the-art multimodal model excelling at a variety of
critical benchmarks for language, text, and image capabilities.

Tiny Aya is a compact 3.35B-parameter multilingual model supporting 70 languages. Its instruction-tuned variants are available on the Cohere API via the
Chat
 endpoint and as open-weight models on
Hugging Face
.

Model Name
Status
Description
Modality
Context Length
Maximum Output Tokens
Endpoints
tiny-aya-global
Live
Tiny Aya Global is a 3.35B instruction-tuned multilingual model with the best balance across languages and regions. Supports 70 languages.
Text
8k
8k
Chat
tiny-aya-earth
Live
Tiny Aya Earth is a 3.35B region-specialized multilingual model, best for West Asian and African languages. Supports 70 languages.
Text
8k
8k
Chat
tiny-aya-fire
Live
Tiny Aya Fire is a 3.35B region-specialized multilingual model, best for South Asian languages. Supports 70 languages.
Text
8k
8k
Chat
tiny-aya-water
Live
Tiny Aya Water is a 3.35B region-specialized multilingual model, best for European and Asia-Pacific languages. Supports 70 languages.
Text
8k
8k
Chat
c4ai-aya-expanse-32b
Live
Aya Expanse is a highly performant 32B multilingual model, designed to rival monolingual performance through innovations in instruction tuning with data arbitrage, preference training, and model merging. Serves 23 languages.
Text
128k
4k
Chat
c4ai-aya-vision-32b
Live
Aya Vision is a state-of-the-art multimodal model excelling at a variety of critical benchmarks for language, text, and image capabilities. Serves 23 languages. This 32 billion parameter variant is focused on state-of-art multilingual performance.
Text, Images
16k
4k
Chat
c4ai-aya-expanse-8b
Retired Apr 4, 2026
Aya Expanse is a highly performant 8B multilingual model, designed to rival monolingual performance through innovations in instruction tuning with data arbitrage, preference training, and model merging. Serves 23 languages.
Text
8k
4k
Chat
c4ai-aya-vision-8b
Retired Apr 4, 2026
Aya Vision is a state-of-the-art multimodal model excelling at a variety of critical benchmarks for language, text, and image capabilities. This 8 billion parameter variant is focused on low latency and best-in-class performance.
Text, Images
16k
4k
Chat
Was this page helpful?
Yes
No
Edit this page
Previous
Cohere Transcribe
Next
Built with