Pricing
- Input Tokens: $0.664 /M tokens
- Output Tokens: $2.738 /M tokens
Input Modalities
- Text
- Vision
Output Modalities
- Text
Frequently asked questions
What is qianfan-ocr-fast?
What is the context length of qianfan-ocr-fast?
How much does qianfan-ocr-fast cost?
What modalities does qianfan-ocr-fast support?
How do I call qianfan-ocr-fast via API?
Who created qianfan-ocr-fast?
More models from Baidu
ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its foundational capabilities and significant improvements in agents, knowledge, reasoning, and deep search. This upgrade uses a decoupled fully-asynchronous reinforcement learning technique to specifically address challenges encountered as large models evolve toward agent-based autonomous decision-making, such as training–inference numerical bias, low utilization of heterogeneous resources, and global issues caused by long-tail effects. It is paired with scaled agent post-training techniques to enhance model capabilities and generalization, enabling a three-step collaboration of environment, expert, and fusion that both ensures training efficiency and significantly improves the model’s stability and performance on complex tasks.
- Input: $ 0.82192 /M
- Output: $ 3.28768 /M
ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities. ERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.
The Ernie-image-Turbo model is an 8-step distilled version of the Ernie-image model, also with 8 billion parameters, offering a 6x speedup compared to pre-distillation, and is suitable for low-latency, local/on-device scenarios.
musesteamer-air-image is a text-to-image model developed by the Baidu Search team aimed at providing extreme cost-effectiveness. It can quickly generate clear images with coherent actions based on user prompts, making it easy to convert users' descriptions into images.
Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on OCR-domain data while retaining appropriate general multimodal capabilities, and it outperforms Qianfan-OCR.
- Input: $ 0.82192 /M
- Output: $ 3.28768 /M
ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities. ERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.
AIHubMix© 2023 - 2026 AIHubMix, LLC