deepseek-ocr

by DeepSeek

DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical character recognition (OCR) and “contextual optical compression.” The model is designed to explore the limits of compressing contextual information from images, efficiently processing documents and converting them into structured text formats such as Markdown. The model requires an image as input.

API Pricing

Input$0.02 / 1M tokens
Output$0.02 / 1M tokens

Specifications

Context window8,000 tokens
Modalitiestext, image

More from DeepSeek

Use deepseek-ocr via the AIHubMix unified API — one interface for every major LLM.