Command A is Cohere most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.
Pricing
- Input Tokens: $2.5 /M tokens
- Output Tokens: $10 /M tokens
- Cache Read: $0 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Context length
- 256K tokens
Max output
- 8K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Try this model
Python