Pricing
- Input Tokens: $2.000 /M tokens
- Output Tokens: $0.000 /M tokens
- Cache Read: $0.000 /M tokens
Image Generation
Input Modalities
Output Modalities
- Image
Frequently asked questions
What is Musesteamer Air Image?
How much does Musesteamer Air Image cost?
How do I call Musesteamer Air Image via API?
Who created Musesteamer Air Image?
More models from Baidu
See all Baidu models →ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its foundational capabilities and significant improvements in agents, knowledge, reasoning, and deep search. This upgrade uses a decoupled fully-asynchronous reinforcement learning technique to specifically address challenges encountered as large models evolve toward agent-based autonomous decision-making, such as training–inference numerical bias, low utilization of heterogeneous resources, and global issues caused by long-tail effects. It is paired with scaled agent post-training techniques to enhance model capabilities and generalization, enabling a three-step collaboration of environment, expert, and fusion that both ensures training efficiency and significantly improves the model’s stability and performance on complex tasks.
- Input: $ 0.82192 /M
- Output: $ 3.28768 /M
ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities. ERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.
The Ernie-image-Turbo model is an 8-step distilled version of the Ernie-image model, also with 8 billion parameters, offering a 6x speedup compared to pre-distillation, and is suitable for low-latency, local/on-device scenarios.
Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on OCR-domain data while retaining appropriate general multimodal capabilities, and it outperforms Qianfan-OCR.
Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on OCR-domain data while retaining appropriate general multimodal capabilities, and it outperforms Qianfan-OCR.
- Input: $ 0.82192 /M
- Output: $ 3.28768 /M
ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities. ERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.
AIHubMix© 2023 - 2026 AIHubMix, LLC