AI Models & Pricing

Explore 825 AI models with transparent per-token pricing — ChatGPT, Claude, Gemini, DeepSeek, Qwen and more, all through one unified API.

auto

by OpenAI

AIHubMix Smart Router: Fill in the model name as auto, and the gateway will automatically…

$2/1M in · $2/1M out
1,000,000 tokens context

gpt-5.6-luna

by OpenAI

GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads. It roughly…

$1/1M in · $6/1M out
1,050,000 tokens context

gpt-5.6-sol

by OpenAI

GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving…

$5/1M in · $30/1M out
1,050,000 tokens context

gpt-5.6-terra

by OpenAI

GPT-5.6 Terra is designed for workloads that balance intelligence and cost. It roughly…

$2.5/1M in · $15/1M out
1,050,000 tokens context

grok-4.5

by Grok

Grok 4.5 was trained on datasets spanning knowledge in coding, science, engineering, and…

$2/1M in · $6/1M out
1,000,000 tokens context

gemini-3.6-flash

by Google

Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…

$1.5/1M in · $7.5/1M out
1,048,576 tokens context

claude-sonnet-5

by Anthropic

Claude Sonnet 5 is the next generation of Anthropic's Sonnet model family. It is a…

$2/1M in · $10/1M out
1,000,000 tokens context

kimi-k3

by Moonshot AI

Kimi K3 is Kimi’s flagship model for long-horizon coding and end-to-end knowledge work…

$3/1M in · $15/1M out
1,048,576 tokens context

qwen3.8-max-preview

by Qwen

Launch offer: Credits are consumed at just 10% of the standard rate, effectively 10× your…

$0.17/1M in · $0.51/1M out
983,616 tokens context

gemini-3.1-flash-lite-image

by Google

Google's newest, most compact, and most cost-effective image generation and editing…

$0.25/1M in · $1.5/1M out

gemini-3.5-flash-lite

by Google

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

gemini-3.5-flash-lite-free

by Google

Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only…

1,048,576 tokens context

gemini-3.6-flash-free

by Google

Gemini 3.6 Flash free version: fFree model resources are limited and provided only for…

1,000,000 tokens context

glm-5.2

by Z.AI

GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable…

$1.13/1M in · $3.94/1M out
1,000,000 tokens context

glm-5.2-fast-preview

by Z.AI

GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2…

$2.25/1M in · $7.89/1M out
1,000,000 tokens context

muse-spark-1.1

by Meta

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It…

$1.38/1M in · $4.67/1M out
1,000,000 tokens context

claude-fable-5

by Anthropic

Anthropic's most capable widely released model, for the most demanding reasoning and…

$11/1M in · $55/1M out
1,000,000 tokens context

claude-opus-4-8

by Anthropic

Claude Opus 4.8 is Anthropic’s newest and most powerful publicly available model. It is…

$5/1M in · $25/1M out
200,000 tokens context

hy3

by Hunyuan

The Hy3 official version is honed for real-world business scenarios, using a…

$0.16/1M in · $0.62/1M out
256,000 tokens context

doubao-seed-2-1-pro

by ByteDance

A new generation of large models moving toward production-grade intelligence…

$0.93/1M in · $4.65/1M out
256,000 tokens context

doubao-seed-2-1-turbo

by ByteDance

Balancing performance and cost, comprehensively upgrading coding, agent, and multimodal…

$0.46/1M in · $2.32/1M out
256,000 tokens context

mai-image-2.5-pro

by Microsoft

MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With…

$5/1M in · $5/1M out

gemini-3.5-flash

by Google

Gemini 3.5 Flash provides sustained frontier-level intelligence optimized for real-world…

$1.5/1M in · $9/1M out
1,000,000 tokens context

grok-build-0.1

by Grok

Fast coding model trained specifically for agentic coding workflows.

$1/1M in · $2/1M out
256,000 tokens context

mai-image-2.5

by Microsoft

MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With…

$5/1M in · $5/1M out

mai-image-2.5-flash

by Microsoft

MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With…

$1.75/1M in · $1.75/1M out

coding-kimi-k3

by Moonshot AI
$0.24/1M in · $0.88/1M out
1,048,576 tokens context

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

coding-glm-5.2-free

by Z.AI

coding-glm-5.2-free is the open and free version of coding-glm-5.2. To ensure stable…

coding-kimi-k3-free

by Moonshot AI

coding-kimi-k3-free is the open and free version of coding-kimi-k3. To ensure stable…

1,048,576 tokens context

gemini-3.1-flash-image

by Google

gemini-3.1-flash-image (Nano Banana 2) features professional-grade visual intelligence…

$0.5/1M in · $3/1M out

gpt-oss-20b-free

by Openai

Developed by OpenAI, gpt-oss-20b-free is an open-weight 21B parameter model released…

131,072 tokens context

kimi-k2.7-code

by Moonshot AI

Kimi K2.7 Code is Kimi’s most intelligent Coding model, capable of completing programming…

$0.95/1M in · $4/1M out
262,144 tokens context

kimi-k2.7-code-highspeed

by Moonshot AI

High-Speed version of Kimi K2.7 Code model, with output speed of approximately 180…

$1.9/1M in · $8/1M out
262,144 tokens context

gemini-3-pro-image

by Google

Gemini-3-Pro-Image (Nano Banana Pro) is a high-performance image generation and editing…

$2/1M in · $12/1M out

gpt-4o-transcribe-diarize

by OpenAI

GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in…

$2.5/1M in · $10/1M out
16,000 tokens context

gpt-audio-1.5

by OpenAI

The gpt-audio model is OpenAI's first officially released (generally available) audio…

$2.5/1M in · $10/1M out
128,000 tokens context

hy-3d-3.1

by Hunyuan

Using the Hunyuan Sheng 3D 3.1 model, it can generate higher-precision and higher-quality…

$2/1M in · $2/1M out

kling-v3-omni

by KLing

VIDEO 3.0 Omni: All-in-One Multimodal Input, Voice-Driven Characters, Direct Audio-Visual…

$2/1M in

kling-video-o1

by KLing

Kling Video O1 is a major unified multimodal video model launched by Kuaishou. It…

$2/1M in

longcat-2.0

by Meituan

Designed for agent development scenarios, it natively supports tool invocation…

$0.77/1M in · $3.1/1M out

nemotron-nano-9b-v2-free

by Nvidia

NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA…

128,000 tokens context

hy3-preview

by Hunyuan

Hunyuan Hy3 preview is designed for agent workloads, adopting a MoE architecture with…

$0.17/1M in · $0.57/1M out
256,000 tokens context

minimax-m3

by Minimax

The MiniMax M3 is a flagship programming model built for real-world productivity. As a…

$0.29/1M in · $1.15/1M out
204,800 tokens context

nemotron-nano-12b-v2-vl-free

by Nvidia

Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open…

128,000 tokens context

qwen3.7-plus

by Qwen

The Qwen 3.7 series' mid-to-high cost-performance "Plus" model builds on strong text…

$0.28/1M in · $1.13/1M out
991,000 tokens context

step-3.7-flash

by StepFun

step-3.7-flash is stepfun's flagship inference model, designed for high-complexity tasks…

$0.22/1M in · $1.32/1M out
256,000 tokens context

claude-opus-4-8-think

by Anthropic

The claude-opus-4-8-think model has adaptive thinking mode pre-enabled; the default…

$5/1M in · $25/1M out
200,000 tokens context

nemotron-3-super-120b-a12b-free

by Nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid…

262,144 tokens context

laguna-m.1-free

by Poolside

Laguna M.1 is the flagship coding agent model from Poolside, optimized for complex…

262,144 tokens context

nemotron-3-nano-omni-30b-a3b-reasoning-free

by Nvidia

Developed by Nvidia, NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…

256,000 tokens context

nemotron-3-ultra-550b-a55b-free

by Nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring…

1,000,000 tokens context

qwen3.7-max

by Qwen

The Max model, the largest and most capable in the Qwen3.7 series, is currently offering…

$1.69/1M in · $5.07/1M out
991,000 tokens context

gpt-image-2

by OpenAI

GPT-image-2 is OpenAI's latest cutting-edge image generation model. Key value adds…

$5/1M in · $30/1M out

nemotron-3.5-content-safety-free

by Nvidia

Developed by NVIDIA, Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal…

128,000 tokens context

coding-glm-5.2

by Z.AI

Currently, the special resources for this model are limited, but due to its popularity…

$0.06/1M in · $0.22/1M out

ernie-5.1

by Baidu

ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its…

$0.56/1M in · $2.54/1M out
119,000 tokens context

gemini-3.1-flash-lite

by Google

gemini-3.1-flash-lite is currently Google's latest and most cost-effective model…

$0.25/1M in · $1.5/1M out
1,000,000 tokens context

gemini-3.1-flash-lite-nothink

by Google

gemini-3.1-flash-lite is currently Google's latest and most cost-effective model…

$0.25/1M in · $1.5/1M out
1,000,000 tokens context

grok-4.3

by Grok

Grok 4.3 is amongst the leading models in intelligence and well priced when comparing to…

$1.25/1M in · $2.5/1M out
1,000,000 tokens context

happyhorse-1.0-i2v

by Qwen

HappyHorse-1.0-I2V supports image-to-video generation, featuring highly faithful dynamic…

$2/1M in

happyhorse-1.0-r2v

by Qwen

HappyHorse-1.0-R2V supports reference-guided video generation, offering more stable…

$2/1M in

happyhorse-1.0-t2v

by Qwen

HappyHorse-1.0-T2V supports text-to-video generation, featuring highly faithful dynamic…

$2/1M in

happyhorse-1.0-video-edit

by Qwen

HappyHorse-1.0-Video-Edit supports video editing, allows editing videos via natural…

$2/1M in

north-mini-code-free

by Cohere

Developed by Cohere, north-mini-code-free is the debut model of the North family and…

256,000 tokens context

gpt-5.5

by OpenAI

GPT-5.5 raises the baseline for complex production workflows. It’s a strong fit for…

$5/1M in · $30/1M out
1,050,000 tokens context

gpt-5.5-pro

by OpenAI

Please note: this model is extremely expensive and very slow. If a request fails due to…

$30/1M in · $180/1M out
1,050,000 tokens context

laguna-xs-2.1-free

by Poolside

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside…

262,144 tokens context

deepseek-v4-flash

by DeepSeek

DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…

$0.15/1M in · $0.31/1M out
1,000,000 tokens context

deepseek-v4-pro

by DeepSeek

DeepSeek-V4 features an ultra-long context of one million characters and achieves leading…

$0.46/1M in · $0.93/1M out
1,000,000 tokens context

gemma-4-31b-it-free

by Google

Gemma 4 31B Instruct is a 30.7B dense multimodal model developed by Google DeepMind that…

262,144 tokens context

command-a-plus-05-2026

by Cohere

Cohere's stronger command model for multilingual agents and enterprise workflows

$2.5/1M in · $10/1M out
128,000 tokens context

doubao-seedream-5.0-pro

by ByteDance

Seedream-5.0-pro is the latest image-creation model released by ByteDance. The model…

$2/1M in

ernie-5.0

by Baidu

ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE…

$0.82/1M in · $3.29/1M out
119,000 tokens context

kimi-k2.6

by Moonshot AI

Kimi K2.6 is Kimi's latest and most intelligent model, with stronger and more stable…

$0.95/1M in · $4/1M out
262,144 tokens context

laguna-s-2.1-free

by Poolside

Laguna S 2.1 is the latest coding agent model from Poolside, featuring an impressive…

262,144 tokens context

qwen3.6-max-preview

by Qwen

The Max model Preview version, the largest and most capable model in the Qwen3.6 series…

$1.27/1M in · $7.61/1M out
240,000 tokens context

xiaomi-mimo-v2.5

by Xiaomi

MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…

$0.15/1M in · $0.31/1M out
256,000 tokens context

xiaomi-mimo-v2.5-pro

by Xiaomi

MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…

$0.48/1M in · $0.96/1M out
1,000,000 tokens context

claude-opus-4-7

by Anthropic

Claude Opus 4.7 is Anthropic’s latest and most powerful publicly available model. It has…

$5/1M in · $25/1M out
200,000 tokens context

claude-opus-4-7-think

by Anthropic

The claude-opus-4-7-think model has adaptive thinking mode pre-enabled; the default…

$5/1M in · $25/1M out
200,000 tokens context

gpt-chat-latest

by OpenAI

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to…

$5/1M in · $30/1M out
1,050,000 tokens context

nemotron-3-nano-30b-a3b-free

by Nvidia

NVIDIA Nemotron 3 Nano 30B A3B is a highly efficient small language Mixture of Experts…

256,000 tokens context

qwen3.6-27b

by Qwen

The Qwen3.6 series 27B native vision-language Dense model. Compared with the 3.5-27B, the…

$0.42/1M in · $2.53/1M out
254,000 tokens context

qwen3.6-35b-a3b

by Qwen

Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding…

$0.25/1M in · $1.52/1M out
254,000 tokens context

qwen3.6-flash

by Qwen

Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding…

$0.17/1M in · $1.01/1M out
991,000 tokens context

cohere-rerank-v4.0-fast

by Cohere

Rerank 4 is the most advanced set of reranker models available today, purpose-built to…

$0.07/1M in

cohere-rerank-v4.0-pro

by Cohere

Rerank 4 is the most advanced set of reranker models available today, purpose-built to…

$0.07/1M in

gemma-4-26b-a4b-it-free

by Google

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model developed by…

262,144 tokens context

grok-4-20-non-reasoning

by Grok

Grok 4.2 is xAI’s latest large language model, built for strong reasoning, multimodal…

$2/1M in · $6/1M out
2,000,000 tokens context

grok-4-20-reasoning

by Grok

Grok 4.2 is xAI’s latest large language model, built for strong reasoning, multimodal…

$2/1M in · $6/1M out
2,000,000 tokens context

mai-image-2e

by Microsoft

MAI-Image-2-Efficient is designed for builders who need high-quality image generation at…

$5/1M in · $19.5/1M out

qwen-image-2.0

by Qwen

The Qwen-Image-2.0 series accelerated models integrate image generation and image…

$2/1M in

qwen-image-2.0-pro

by Qwen

The Qwen-Image-2.0 full-powered models achieve the integration of image generation and…

$2/1M in

coding-minimax-m3-free

by Minimax

coding-minimax-m3-free is a free and open version offered by AIHubMix specifically for…

204,800 tokens context

doubao-seedance-2-0-260128

by Doubao

The Doubao large-model team has launched a new-generation professional-grade multimodal…

$2/1M in

doubao-seedance-2-0-fast-260128

by Doubao

Seedance 2.0 fast is a next-generation multimodal video-creation model launched by the…

$2/1M in

doubao-seedance-2-0-mini-260615

by ByteDance

Seedance 2.0 mini is a next-generation, cost-effective video generation model launched to…

$2/1M in

glm-5.1

by Z.AI

GLM-5.1 is Zhipu's latest flagship model, with greatly enhanced coding capabilities and…

$0.84/1M in · $3.38/1M out
200,000 tokens context

glm-image

by Z.AI

GLM-Image is Zhipu AI's new flagship image generation model. The model is trained…

$2/1M in · $2/1M out

qwen3.6-plus

by Qwen

Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding…

$0.28/1M in · $1.69/1M out
991,000 tokens context

wan2.7-i2v

by Qwen

Wanxiang 2.7 — image-to-video: performance capabilities comprehensively upgraded…

$2/1M in

wan2.7-r2v

by Qwen

Wanxiang 2.7 — reference-driven video generation: more stable references for characters…

$2/1M in

wan2.7-t2v

by Qwen

Wanxiang 2.7 — text-to-video: performance capabilities comprehensively upgraded…

$2/1M in

wan2.7-videoedit

by Qwen

Wanxiang 2.7 — video editing: edit videos using natural-language commands, supporting…

$2/1M in

cc-k2.6-code-preview

by Moonshot AI

for claude code

$0.2/1M in · $0.2/1M out

gemma-4-26b-a4b-it

by Google

A Mixture-of-Experts model that activates only 4B parameters per inference,delivering…

$0.14/1M in · $0.4/1M out
262,100 tokens context

gemma-4-31b-it

by Google

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text…

$0.14/1M in · $0.4/1M out
262,100 tokens context

gpt-5.4

by OpenAI

GPT-5.4 is our frontier model for complex professional work.Reasoning.effort supports…

$2.5/1M in · $15/1M out
400,000 tokens context

wan2.7-image

by Qwen

Wanxiang 2.7 — image generation and editing: supports text-to-image, text-to-multi-image…

$2/1M in · $2/1M out

wan2.7-image-pro

by Qwen

Wanxiang 2.7 — image generation and editing: supports text-to-image, text-to-multi-image…

$2/1M in · $2/1M out

claude-sonnet-4-6

by Anthropic

Claude Sonnet 4.6 delivers frontier intelligence at scale—built for coding, agents, and…

$3/1M in · $15/1M out
1,000,000 tokens context

coding-xiaomi-mimo-v2.5

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.16/1M out

coding-xiaomi-mimo-v2.5-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.4/1M out

doubao-seed-2-0-lite-260428

by Doubao

Doubao Coding model optimized for real-world programming environments that can reliably…

$0.09/1M in · $0.54/1M out
256,000 tokens context

doubao-seed-2-0-mini-260428

by Doubao

Doubao Coding model optimized for real-world programming environments that can reliably…

$0.03/1M in · $0.28/1M out
256,000 tokens context

gemini-3.1-flash-image-preview

by Google

gemini-3.1-flash-image-preview (Nano Banana 2) features professional-grade visual…

$0.5/1M in · $3/1M out

gemini-3.1-flash-lite-preview

by Google

gemini-3.1-flash-lite-preview is currently Google's latest and most cost-effective model…

$0.25/1M in · $1.5/1M out
1,000,000 tokens context

gemini-3.1-pro-preview

by Google

Gemini 3.1 Pro Preview is designed to further optimize the performance and reliability of…

$2/1M in · $12/1M out
1,000,000 tokens context

gemini-3.1-pro-preview-customtools

by Google

gemini-3.1-pro-preview-customtools For users who build applications mixing bash and…

$2/1M in · $12/1M out
1,000,000 tokens context

gemini-3.1-pro-preview-search

by Google

Gemini-3.1-pro-preview-search integrates Google's official search functionality; the…

$2/1M in · $12/1M out

gpt-5.4-mini

by OpenAI

GPT-5.4 mini is a faster, more efficient model that inherits the advantages of GPT-5.4…

$0.75/1M in · $4.5/1M out
400,000 tokens context

gpt-5.4-nano

by OpenAI

GPT-5.4 nano is designed for tasks where speed and cost are most important, such as…

$0.2/1M in · $1.25/1M out
400,000 tokens context

gpt-5.5-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

1,050,000 tokens context

qwen3.5-plus

by Qwen

The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that…

$0.11/1M in · $0.66/1M out
991,000 tokens context

claude-sonnet-4-6-think

by Anthropic

Claude sonnet 4.6 does not enable reasoning mode by default. To access its deep reasoning…

$3/1M in · $15/1M out
200,000 tokens context

coding-xiaomi-mimo-v2-omni

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.4/1M out

coding-xiaomi-mimo-v2-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.6/1M out

gpt-5.3-chat-latest

by OpenAI

GPT-5.3Chat refers to the GPT-5.3 snapshot currently used in ChatGPT and is optimized for…

$1.75/1M in · $14/1M out
128,000 tokens context

gpt-5.3-codex

by OpenAI

GPT-5.3-Codex is optimized for agentic coding tasks in Codex or similar environments…

$1.75/1M in · $14/1M out
400,000 tokens context

gpt-image-2-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

qwen3.5-122b-a10b

by Qwen

The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that…

$0.11/1M in · $0.9/1M out
991,000 tokens context

qwen3.5-27b

by Qwen

The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that…

$0.08/1M in · $0.68/1M out
991,000 tokens context

qwen3.5-35b-a3b

by Qwen

The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that…

$0.06/1M in · $0.45/1M out
991,000 tokens context

qwen3.5-397b-a17b

by Qwen

The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that…

$0.16/1M in · $0.99/1M out
991,000 tokens context

qwen3.5-flash

by Qwen

The Qwen3.5 native vision-language Flash series models are designed with a hybrid…

$0.03/1M in · $0.28/1M out
991,000 tokens context

coding-glm-5.1

by Z.AI

Only supports OpenAI-compatible formats.

$0.06/1M in · $0.22/1M out

doubao-seed-2-0-pro

by Doubao

Doubao flagship all-purpose general model, targeting complex reasoning and long-chain…

$0.48/1M in · $2.41/1M out
256,000 tokens context

gpt-5.4-high

by OpenAI

GPT-5.4 supports configurable reasoning effort only through the /responses endpoint. To…

$2.5/1M in · $15/1M out
400,000 tokens context

gpt-5.4-low

by OpenAI

GPT-5.4 supports configuring reasoning strength only through the /responses endpoint. To…

$2.5/1M in · $15/1M out
400,000 tokens context

gpt-5.4-pro

by OpenAI

Please note: this model is extremely expensive and very slow. If a request fails due to…

$30/1M in · $180/1M out
1,050,000 tokens context

qwen3-coder-next

by Qwen

The Qwen3 series is a next-generation code-generation model with results close to…

$0.14/1M in · $0.55/1M out
2,000,000 tokens context

xiaomi-mimo-v2-omni-free

by Xiaomi

xiaomi-mimo-v2-omni-free is the open free version of xiaomi-mimo-v2-omni. To ensure…

256,000 tokens context

xiaomi-mimo-v2-pro-free

by Xiaomi

xiaomi-mimo-v2-pro-free is the open free version of xiaomi-mimo-v2-pro. To ensure stable…

256,000 tokens context

xiaomi-mimo-v2.5-free

by Xiaomi

xiaomi-mimo-v2.5-free is the open free version of xiaomi-mimo-v2.5. To ensure stable…

256,000 tokens context

xiaomi-mimo-v2.5-pro-free

by Xiaomi

xiaomi-mimo-v2.5-pro-free is the open free version of xiaomi-mimo-v2.5-pro5. To ensure…

256,000 tokens context

claude-opus-4-6

by Anthropic

Claude Opus 4.6 is Anthropic’s latest state-of-the-art reasoning model. It features an…

$5/1M in · $25/1M out
200,000 tokens context

coding-glm-5.1-free

by Z.AI

coding-glm-5.1-free is the open and free version of coding-glm-5.1. To ensure stable…

coding-minimax-m2.7-free

by Minimax

coding-minimax-m2.7-free is a free and open version offered by AIHubMix specifically for…

204,800 tokens context

glm-5

by Z.AI

GLM-5 is an advanced, open-source large language model designed for developers tackling…

$0.88/1M in · $2.82/1M out
202,752 tokens context

glm-5v-turbo

by Z.AI

GLM-5V-Turbo is Zhipu's first multimodal coding foundation model, built for visual…

$0.7/1M in · $3.1/1M out
200,000 tokens context

minimax-m2.7

by Minimax

MiniMax M2.7 can autonomously build complex Agent Harnesses and, leveraging capabilities…

$0.3/1M in · $1.18/1M out
200,000 tokens context

claude-opus-4-6-think

by Anthropic

Claude Opus 4.6 does not enable reasoning mode by default. To access its deep reasoning…

$5/1M in · $25/1M out
200,000 tokens context

coding-glm-5-free

by Z.AI

coding-glm-5-free is the open and free version of coding-glm-5. To ensure stable service…

coding-glm-5-turbo-free

by Z.AI

coding-glm-5-turbo-free is the open and free version of coding-glm-5-turbo. To ensure…

coding-minimax-m2.5-free

by Minimax

coding-minimax-m2.5-free is a free and open version offered by AIHubMix specifically for…

204,800 tokens context

doubao-seed-2-0-code-preview

by Doubao

The Doubao 2.0 series is a coding model optimized for real programming environments…

$0.48/1M in · $2.41/1M out
256,000 tokens context

doubao-seed-2-0-lite-260215

by Doubao

Doubao Coding model optimized for real-world programming environments that can reliably…

$0.09/1M in · $0.54/1M out
256,000 tokens context

doubao-seed-2-0-mini

by Doubao

Doubao 2.0 series is designed for low-latency, high-concurrency, and cost-sensitive…

$0.03/1M in · $0.3/1M out
256,000 tokens context

gemini-3-flash-preview

by Google

gemini-3-flash-preview is Google's latest released, most balanced model, excelling in…

$0.5/1M in · $3/1M out
1,048,576 tokens context

gemini-3-flash-preview-search

by Google

Gemini-3-flash-preview-search integrates Google's official search functionality; the…

$0.5/1M in · $3/1M out
1,048,576 tokens context

glm-5-turbo

by Z.AI

GLM-5-Turbo is a foundational model deeply optimized for the OpenClaw scenario. From the…

$1.2/1M in · $4/1M out
202,752 tokens context

cc-glm-5.1

by Z.AI

Supports Claude native interface, can be directly requested in Claude Code.

$0.06/1M in · $0.22/1M out

claude-opus-4-5

by Anthropic

Claude Opus 4.5 is Anthropic’s latest frontier reasoning model, optimized for complex…

$5/1M in · $25/1M out
200,000 tokens context

claude-opus-4-5-think

by Anthropic

Claude Opus 4.5 does not enable reasoning mode by default. To access its deep reasoning…

$5/1M in · $25/1M out
200,000 tokens context

embed-v-4-0

by Cohere

Cohere’s Embed 4 is a multilingual multimodal embedding model. It is capable of…

$0.12/1M in
128,000 tokens context

ernie-image-turbo

by Baidu

The Ernie-image-Turbo model is an 8-step distilled version of the Ernie-image model, also…

$2/1M in

gemini-3.1-flash-image-preview-free

by Google

This model is the free trial version of gemini-3.1-flash-image-preview (officially…

mimo-v2-omni

by Xiaomi

MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution…

$0.44/1M in · $2.2/1M out
256,000 tokens context

mimo-v2-pro

by Xiaomi

Xiaomi MiMo-V2-Pro is built for high-intensity agent work scenarios in the real world. It…

$1.1/1M in · $3.3/1M out
1,000,000 tokens context

cohere-command-a

by Cohere

Command A is Cohere most performant model to date, excelling at tool use, agents…

$2.5/1M in · $10/1M out

gemini-3-flash-preview-free

by Google

gemini-3-flash-preview-free is the free, publicly available version of…

1,048,576 tokens context

cc-minimax-m3

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

coding-minimax-m3

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

gpt-4.1-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

1,047,576 tokens context

gpt-4.1-mini-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

1,047,576 tokens context

gpt-4.1-nano-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

1,047,576 tokens context

gpt-4o-free

by OpenAI

This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the…

1,047,576 tokens context

coding-glm-5

by Z.AI

Only supports OpenAI-compatible formats.

$0.06/1M in · $0.22/1M out

coding-glm-5-turbo

by Z.AI

Only supports OpenAI-compatible formats.

$0.06/1M in · $0.22/1M out

glm-4.7

by Z.AI

GLM-4.7 is Zhiyuan's latest flagship model. GLM-4.7 enhances coding capabilities…

$0.27/1M in · $1.1/1M out
200,000 tokens context

veo-3.1-lite-generate-preview

by Google

Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p…

$2/1M in

glm-4.7-flash-free

by Z.AI

The glm-4.7-flash free model has usage restrictions to ensure stable service operation: a…

coding-glm-4.7-free

by Z.AI

coding-glm-4.7-free is the open and free version of coding-glm-4.7. To ensure stable…

doubao-seedance-1-5-pro-251215

by Doubao

The Doubao video generation model Seedance 1.5 Pro, as a world-leading video generation…

$2/1M in

doubao-seedance-1-0-pro-250528

by Doubao

Seedance 1.0 Pro is a foundational video-generation model that supports multi-shot…

$2/1M in

doubao-seedance-1-0-pro-fast-251015

by Doubao

Seedance 1.0 Pro Fast is a comprehensive model that delivers rock-bottom prices and peak…

$2/1M in

gemini-3-pro-image-preview

by Google

Gemini-3-Pro-Image-Preview (Nano Banana Pro) is a high-performance image generation and…

$2/1M in · $12/1M out

gemini-embedding-2

by Google

Google's first multimodal embedding model .

$0.2/1M in

deepinfra-gemma-4-26b-a4b-it

by Google

A Mixture-of-Experts model that activates only 4B parameters per inference,delivering…

$0.09/1M in · $0.38/1M out
262,100 tokens context

gpt-5.2-codex

by OpenAI

GPT-5.2-Codex is an upgraded version of GPT-5.2, optimized for agentic coding tasks in…

$1.75/1M in · $14/1M out
400,000 tokens context

doubao-seedream-5.0-lite

by Doubao

Doubao-Seedream-5.0-lite is the latest image-creation model released by ByteDance. For…

$2/1M in

gpt-image-1.5

by OpenAI

GPT Image 1.5 is a new image generation model powered by OpenAI’s flagship visual…

$5/1M in · $10/1M out

gemini-3.1-flash-lite-preview-nothink

by Google

gemini-3.1-flash-lite-preview is currently Google's latest and most cost-effective model…

$0.25/1M in · $1.5/1M out
1,000,000 tokens context

gpt-5.2

by OpenAI

GPT-5.2 is an advanced general-purpose model that improves on GPT-5.1 with more reliable…

$1.75/1M in · $14/1M out
400,000 tokens context

gpt-5.2-chat-latest

by OpenAI

GPT-5.2Chat refers to the GPT-5.2 snapshot currently used in ChatGPT and is optimized for…

$1.75/1M in · $14/1M out
128,000 tokens context

gpt-5.2-high

by OpenAI

GPT-5.2 supports configurable reasoning effort only through the /responses endpoint. To…

$1.75/1M in · $14/1M out
400,000 tokens context

gpt-5.2-low

by OpenAI

GPT-5.2 supports configuring reasoning strength only through the /responses endpoint. To…

$1.75/1M in · $14/1M out
400,000 tokens context

gpt-5.2-pro

by OpenAI

GPT-5.2 pro is available in the Responses API only to enable support for multi-turn model…

$21/1M in · $168/1M out
400,000 tokens context

gpt-5.1

by OpenAI

GPT-5 is OpenAI’s most advanced language model, designed for complex tasks that require…

$1.25/1M in · $10/1M out
400,000 tokens context

gpt-5.1-codex-max

by OpenAI

GPT-5.1-Codex-Max is a frontier programming model built for the agent-driven era. Powered…

$1.25/1M in · $10/1M out
400,000 tokens context

doubao-seed-1-8

by Doubao

Doubao's strongest multimodal Agent model Seed1.8 has powerful multimodal capabilities…

$0.11/1M in · $0.27/1M out
256,000 tokens context

gpt-5.1-chat-latest

by OpenAI

GPT-5.1 Chat refers to the GPT-5.1 snapshot currently used in ChatGPT and is optimized…

$1.25/1M in · $10/1M out
128,000 tokens context

gpt-5.1-codex

by OpenAI

GPT-5.1-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or…

$1.25/1M in · $10/1M out
400,000 tokens context

gpt-5.1-codex-mini

by OpenAI

GPT-5.1 Codex mini is a smaller, more cost-effective, less-capable version of…

$0.25/1M in · $2/1M out
400,000 tokens context

claude-haiku-4-5

by Anthropic

Claude Haiku 4.5 is a fast, affordable, and highly capable AI model, excelling at coding…

$1.1/1M in · $5.5/1M out
204,800 tokens context

claude-sonnet-4-5

by Anthropic

Sonnet 4.5 is the best model in the world for agents, coding, and computer usage. It is…

$3.3/1M in · $16.5/1M out
1,000,000 tokens context

claude-sonnet-4-5-think

by Anthropic

Claude Sonnet 4.5 does not enable reasoning mode by default. To access its deep reasoning…

$3.3/1M in · $16.5/1M out
1,000,000 tokens context

grok-4.20-multi-agent-0309

by Grok

Grok 4.20 is our newest flagship model with industry-leading speed and agentic tool…

$2/1M in · $6/1M out
2,000,000 tokens context

mistral-large-3

by Mistral

Mistral Large 3 is a MoE model with 67.5B total parameters and 41B active parameters…

$0.5/1M in · $1.5/1M out
256,000 tokens context

cc-glm-5

by Z.AI

Supports Claude native interface, can be directly requested in Claude Code.

$0.06/1M in · $0.22/1M out

cc-glm-5-turbo

by Z.AI

Supports Claude native interface, can be directly requested in Claude Code.

$0.06/1M in · $0.22/1M out

cloudflare-glm-5.2

by Z.AI
$1.4/1M in · $4.4/1M out

gemini-2.5-flash-image

by Google

Gemini 2.5 Flash Image (Nano-Banana) is a state-of-the-art image generation and editing…

$0.3/1M in · $2.5/1M out
32,800 tokens context

grok-4-1-fast-non-reasoning

by Grok

Grok 4.1 is a new conversational model with significant improvements in real-world…

$0.2/1M in · $0.5/1M out
2,000,000 tokens context

grok-4-1-fast-reasoning

by Grok

Grok 4.1 is a new conversational model with significant improvements in real-world…

$0.2/1M in · $0.5/1M out
2,000,000 tokens context

grok-code-fast-1

by Grok

Grok 4.1 is a new conversational model with significant improvements in real-world…

$0.2/1M in · $0.5/1M out
256,000 tokens context

k2.6-code-preview-free

by Moonshot AI

kimi-for-coding-free is a free and open version offered by AIHubMix specifically for Kimi…

256,000 tokens context

mimo-v2-flash

by Xiaomi

MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion…

$0.19/1M in · $0.58/1M out

musesteamer-air-image

by Baidu

musesteamer-air-image is a text-to-image model developed by the Baidu Search team aimed…

$2/1M in

qwen3.6-plus-preview-free

by Qwen

This model has been removed from the platform.

1,000,000 tokens context

zai-glm-5-turbo

by Z.AI
$1.2/1M in · $4/1M out

gpt-5

by OpenAI

GPT-5 is OpenAI’s most advanced general-purpose model, delivering major improvements in…

$1.25/1M in · $10/1M out
400,000 tokens context

deepseek-v3.2

by DeepSeek

DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…

$0.3/1M in · $0.45/1M out
128,000 tokens context

deepseek-v3.2-think

by DeepSeek

DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse…

$0.3/1M in · $0.45/1M out
128,000 tokens context

gpt-5-codex

by OpenAI

GPT-5-Codex is a version of GPT-5 optimized for autonomous coding tasks in Codex or…

$1.25/1M in · $10/1M out
400,000 tokens context

DeepSeek-V3.1-Terminus

by DeepSeek

DeepSeek-V3.1 non-thinking mode has now been updated to the DeepSeek-V3.1-Terminus…

$0.56/1M in · $1.68/1M out
160,000 tokens context

DeepSeek-V3.1-Think

by DeepSeek

Thinking mode of DeepSeek-V3.1; DeepSeek V3.1 is a text generation model provided by…

$0.56/1M in · $1.68/1M out
128,000 tokens context

gpt-5-pro

by OpenAI

GPT-5 pro uses more compute to think harder and provide consistently better…

$15/1M in · $120/1M out
400,000 tokens context

gpt-5-mini

by OpenAI

GPT-5 mini is a faster, more cost-efficient version of GPT-5. It's great for well-defined…

$0.25/1M in · $2/1M out
400,000 tokens context

gpt-5-nano

by OpenAI

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, designed specifically…

$0.05/1M in · $0.4/1M out
400,000 tokens context

gpt-5-chat-latest

by OpenAI

GPT-5 Chat points to the GPT-5 snapshot currently used in ChatGPT. GPT-5 is our…

$1.25/1M in · $10/1M out
400,000 tokens context

claude-opus-4-1

by Anthropic

Opus 4.1 is an upgraded version of Claude Opus 4, with improvements mainly in agent…

$16.5/1M in · $82.5/1M out
200,000 tokens context

o3-deep-research

by OpenAI

Only supported through requests to the v1/responses interface. o3-deep-research is OpenAI…

$10/1M in · $40/1M out

kimi-k2.5

by Moonshot AI

Kimi K2.5 is the smartest model of Kimi to date, achieving open-source state-of-the-art…

$0.6/1M in · $3/1M out
256,000 tokens context

qwen3-max-2026-01-23

by Qwen

The snapshot version of the Tongyi Qianwen 3 series Max model is from January 23, 2026…

$0.45/1M in · $1.8/1M out
252,000 tokens context

qwen3-vl-flash

by Qwen

The Qwen3 series of compact visual-understanding models achieves an effective fusion of…

$0.02/1M in · $0.21/1M out
254,000 tokens context

qwen3-vl-flash-2026-01-22

by Qwen

The Qwen3 series of compact visual-understanding models achieves an effective fusion of…

$0.02/1M in · $0.21/1M out
254,000 tokens context

qwen3-vl-plus

by Qwen

The Qwen3 series visual understanding model achieves an effective fusion of thinking and…

$0.14/1M in · $1.37/1M out
256,000 tokens context

cc-minimax-m2.7

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

cc-minimax-m2.7-highspeed

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

minimax-m2.5

by Minimax

The MiniMax M2.5 is a flagship programming model built for real-world productivity. As a…

$0.29/1M in · $1.15/1M out
204,800 tokens context

minimax-m2.5-highspeed

by Minimax

• Same performance as minimax-m2.5 • Significantly faster inference

$0.29/1M in · $1.15/1M out
204,800 tokens context

mm-minimax-m2.7-highspeed

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

coding-minimax-m2.7

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

coding-minimax-m2.7-highspeed

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

cc-minimax-m2.5

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

cc-minimax-m2.5-highspeed

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

coding-minimax-m2.5

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

coding-minimax-m2.5-highspeed

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

doubao-seedream-4-5

by Doubao

Seedream 4.5 is ByteDance's latest multimodal image model, integrating capabilities such…

$2/1M in

sora-2

by OpenAI

Sora-2 is the next-generation text-to-video model evolved from Sora, optimized for higher…

$2/1M in · $2/1M out

sora-2-pro

by OpenAI

OpenAI video model Sora2-pro official API.

$2/1M in · $2/1M out

cc-glm-4.7

by Z.AI

Supports Claude native interface, can be directly requested in Claude Code.

$0.06/1M in · $0.22/1M out

cc-minimax-m2.1

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

coding-glm-4.7

by Z.AI

Only supports OpenAI-compatible formats.

$0.06/1M in · $0.22/1M out

coding-minimax-m2.1

by Minimax
$0.2/1M in · $0.2/1M out
204,800 tokens context

coding-minimax-m2.1-free

by Minimax

coding-minimax-m2.1-free is a free and open version offered by AIHubMix specifically for…

204,800 tokens context

gpt-4o-audio-preview

by OpenAI

OpenAI voice input and output model, with prices consistent with the official ones. For…

$2.5/1M in · $10/1M out
128,000 tokens context

gpt-4o-mini-audio-preview

by OpenAI

openai声音输入输出模型,价格和官方一致,暂时只展示文字部分价格,声音价格见openai官网;后台扣费和官方一致

$0.15/1M in · $0.6/1M out

minimax-m2.1

by Minimax

MiniMax-M2.1 redefines efficiency for intelligent agents. It is a compact, fast, and…

$0.29/1M in · $1.15/1M out
204,800 tokens context

o3

by OpenAI

OpenAI o3 is a powerful model across multiple domains, setting a new standard for coding…

$2/1M in · $8/1M out
200,000 tokens context

wan2.6-i2v

by Qwen

Wan 2.6 - Text-to-Video generation features intelligent storyboard scheduling supporting…

$2/1M in

wan2.6-t2v

by Qwen

Wan 2.6 - Text-to-Video generation features intelligent storyboard scheduling supporting…

$2/1M in

cc-glm-4.6

by Z.AI

for claude code

$0.06/1M in · $0.22/1M out

coding-glm-4.6

by Z.AI
$0.06/1M in · $0.22/1M out

coding-glm-4.6-free

by Z.AI

coding-glm-4.6-free is the open and free version of coding-glm-4.6. To ensure stable…

200,000 tokens context

coding-minimax-m2

by Minimax

coding-minimax-m2 is a free and open version offered by AIHubMix specifically for MiniMax…

$0.2/1M in · $0.2/1M out
204,800 tokens context

coding-minimax-m2-free

by Minimax

coding-minimax-m2-free is a free and open version offered by AIHubMix specifically for…

204,800 tokens context

flux-2-flex

by Flux

FLUX.2 is purpose-built for real-world creative production workflows. It delivers…

$2/1M in

flux-2-pro

by Flux

FLUX.2 is purpose-built for real-world creative production workflows. It delivers…

$2/1M in

gemini-2.5-pro

by Google

Gemini 2.5 Pro is an advanced reasoning model developed by Google, optimized for solving…

$1.25/1M in · $10/1M out
1,048,576 tokens context

glm-4.6

by Z.AI

GLM-4.6 is Zhipu’s latest flagship model (total parameters 355B, activation parameters…

$0.27/1M in · $1.1/1M out
204,800 tokens context

glm-4.6v

by Z.AI

Zhipu's latest visual reasoning model achieves state-of-the-art visual understanding…

$0.14/1M in · $0.41/1M out
128,000 tokens context

glm-ocr

by Z.AI

GLM-OCR is a lightweight professional OCR model with only 0.9B parameters, yet multiple…

$0.03/1M in · $0.03/1M out
32,000 tokens context

kimi-for-coding-free

by Moonshot AI

kimi-for-coding-free is a free and open version offered by AIHubMix specifically for Kimi…

256,000 tokens context

o3-pro

by OpenAI

o3-pro This model only supports Requests API interface requests.The model's thinking time…

$20/1M in · $80/1M out
200,000 tokens context

qianfan-ocr

by Baidu

Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…

$0.06/1M in · $0.25/1M out
32,000 tokens context

qianfan-ocr-fast

by Baidu

Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…

$0.66/1M in · $2.74/1M out
32,000 tokens context

step-3.5-flash

by StepFun

step-3.5-flash is stepfun's flagship inference model, designed for high-complexity tasks…

$0.11/1M in · $0.33/1M out
256,000 tokens context

wan2.2-i2v-plus

by Qwen

The newly upgraded Tongyi Wanxiang 2.2 text-to-video offers higher video quality. It…

$2/1M in

wan2.5-i2v-preview

by Qwen

Tongyi Wanxiang 2.5 - Text-to-Video Preview features a newly upgraded technical…

$2/1M in

wan2.5-t2v-preview

by Qwen

Tongyi Wanxiang 2.5 - Text-to-Video Preview, newly upgraded model architecture, supports…

$2/1M in

gemini-2.5-pro-search

by Google

gemini-2.5-pro-search integrates Google's official search functionality; the search…

$1.25/1M in · $10/1M out
1,048,576 tokens context

kimi-k2-thinking

by Moonshot AI

Kimi K2 Thinking is Moonshot AI's most advanced open-source inference model to date…

$0.55/1M in · $2.19/1M out
262,144 tokens context

gemini-2.5-flash

by Google

Gemini 2.5 Flash is Google’s best model in terms of both performance and cost efficiency…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

gemini-2.5-flash-preview-09-2025

by Google

This latest 2.5 Flash model comes with improvements in two key areas we heard consistent…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

glm-4.5v

by Z.AI

GLM-4.5V is a vision-language foundational model designed for multimodal agent…

$0.27/1M in · $0.82/1M out
64,000 tokens context

gemini-2.5-flash-lite

by Google

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that…

$0.1/1M in · $0.4/1M out
1,048,576 tokens context

gemini-2.5-flash-lite-nothink

by Google

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that…

$0.1/1M in · $0.4/1M out
1,048,576 tokens context

gemini-2.5-flash-lite-preview-09-2025

by Google

gemini-2.5-flash-lite latest preview version

$0.1/1M in · $0.4/1M out
1,048,576 tokens context

gemini-2.5-flash-lite-preview-09-2025-nothink

by Google

gemini-2.5-flash-lite latest preview version

$0.1/1M in · $0.4/1M out
1,048,576 tokens context

gemini-2.5-flash-nothink

by Google

Gemini-2.5-flash defaults to thinking enabled; to disable thinking, request the name…

$0.3/1M in · $2.5/1M out
1,047,576 tokens context

gemini-2.5-flash-search

by Google

gemini-2.5-flash-search integrates Google's official search functionality; the search…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

gemini-2.5-flash-preview-05-20-nothink

by Google

Gemini-2.5-flash-preview-05-20 is enabled by default for thinking; to disable it, request…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

gemini-2.5-flash-preview-05-20-search

by Google

Gemini-2.5 Flash Preview 05-20 Search integrates Google's official search functionality…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

DeepSeek-V3-Fast

by DeepSeek

V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to…

$0.56/1M in · $2.24/1M out
32,000 tokens context

imagen-4.0

by Google

Imagen 4 is a high-quality text-to-image model developed by Google, designed for strong…

$2/1M in · $2/1M out

imagen-4.0-fast-generate-001

by Google

Imagen 4 is a new-generation image generation model designed to balance high-quality…

$2/1M in · $2/1M out

imagen-4.0-generate-001

by Google

Imagen 4 is a new-generation image generation model designed to balance high-quality…

$2/1M in · $2/1M out

imagen-4.0-ultra-generate-001

by Google

Imagen-4.0 最新正式版本

$2/1M in · $2/1M out

imagen-4.0-ultra

by Google

Imagen-4.0 Ultra 最新预览版

$2/1M in · $2/1M out

gpt-image-1

by OpenAI

Azure OpenAI’s gpt-image-1 image generation API offers both text-to-image generation and…

$5/1M in · $40/1M out

gpt-image-1-mini

by OpenAI

OpenAI image generation model gpt-image-1-mini Before use, please run pip install -U…

$5/1M in · $40/1M out

o4-mini

by OpenAI

o4-mini is a remarkably smart model for its speed and cost-efficiency. This allows it to…

$1.1/1M in · $4.4/1M out
200,000 tokens context

DeepSeek-OCR

by DeepSeek

DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical…

$0.02/1M in · $0.02/1M out
8,000 tokens context

alicloud-kimi-k2-instruct

by Moonshot AI

Kimi-K2 is a MoE architecture foundational model with extremely powerful coding and agent…

$0.55/1M in · $2.19/1M out

deepseek-ocr

by DeepSeek

DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical…

$0.02/1M in · $0.02/1M out
8,000 tokens context

ernie-5.0-thinking-exp

by Baidu

ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE…

$0.82/1M in · $3.29/1M out
119,000 tokens context

flux-kontext-max

by Flux
$2/1M in

gemini-2.5-flash-image-preview

by Google

Aihubmix supports the gemini-2.5-flash-image-preview model; you can add extra parameters…

$0.3/1M in · $1.2/1M out
32,800 tokens context

glm-4.5

by ChatGLM

GLM-4.5

$0.4/1M in · $1.6/1M out
131,072 tokens context

gpt-4.1

by OpenAI

The latest flagship multimodal model supports million-token context, with encoding…

$2/1M in · $8/1M out
1,047,576 tokens context

grok-4

by Grok

Grok, their latest and greatest flagship model, offers unparalleled performance in…

$3.3/1M in · $16.5/1M out
256,000 tokens context

grok-4-fast-non-reasoning

by Grok

Grok-4-fast is a cost-effective inference model developed by xAI that delivers…

$0.2/1M in · $0.5/1M out
2,000,000 tokens context

grok-4-fast-reasoning

by Grok

Grok-4-fast is a cost-effective inference model developed by xAI that delivers…

$0.2/1M in · $0.5/1M out
2,000,000 tokens context

kimi-k2-0711

by Moonshot AI

Kimi-K2 is a MoE architecture foundational model with extremely powerful coding and agent…

$0.54/1M in · $2.16/1M out
131,000 tokens context

kimi-k2-instruct

by Moonshot AI

Kimi-K2 is a MoE architecture foundational model with extremely powerful coding and agent…

$0.54/1M in · $2.16/1M out

kimi-k2-turbo-preview

by Moonshot AI

The kimi-k2-turbo-preview model is a high-speed version of kimi-k2, with the same model…

$1.2/1M in · $4.8/1M out
262,144 tokens context

paddleocr-vl-0.9b

by Baidu

PaddleOCR-VL is an advanced and efficient document parsing model specifically designed…

$2/1M in

pp-structurev3

by Baidu

PP-StructureV3 is an efficient and comprehensive document parsing solution that can…

$2/1M in

qwen3-vl-235b-a22b-instruct

by Qwen

The Qwen3 series open-source models include hybrid models, thinking models, and…

$0.27/1M in · $1.1/1M out
131,000 tokens context

qwen3-vl-235b-a22b-thinking

by Qwen

The Qwen3 series open-source models include hybrid models, thinking models, and…

$0.27/1M in · $2.74/1M out
131,000 tokens context

qwen3-vl-30b-a3b-instruct

by Qwen

The Qwen3-VL series’ second-largest MoE model Instruct version offers fast response speed…

$0.1/1M in · $0.41/1M out
128,000 tokens context

qwen3-vl-30b-a3b-thinking

by Qwen

The Qwen3-VL series’ second-largest MoE model Thinking version offers fast response…

$0.1/1M in · $1.03/1M out
128,000 tokens context

veo-3.0-generate-preview

by Google

Veo 3.0 Generate Preview is an advanced AI video generation model that supports…

$2/1M in · $2/1M out

veo-3.1-fast-generate-preview

by Google

Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p…

$2/1M in

veo-3.1-generate-preview

by Google

Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p…

$2/1M in · $2/1M out

aihubmix-router

by OpenAI

New model routing capability; request aihubmix-router to automatically route models based…

$0.4/1M in · $1.6/1M out

gpt-4.1-mini

by OpenAI

Lightweight, high-performance model with million-token context and near-flagship-level…

$0.4/1M in · $1.6/1M out
1,047,576 tokens context

gpt-4.1-nano

by OpenAI

Ultra-lightweight model with million-token context, optimized for speed and low latency…

$0.1/1M in · $0.4/1M out
1,047,576 tokens context

gemini-2.5-pro-preview-05-06

by Google

gemini-2.5-pro latest model

$1.25/1M in · $10/1M out
1,048,576 tokens context

gemini-2.5-pro-preview-03-25

by Google

Supports high concurrency. The Gemini 2.5 Pro preview version is here, with higher…

$1.25/1M in · $10/1M out

gemini-2.5-pro-preview-05-06-search

by Google

Integrated with Google's official search function.

$1.25/1M in · $10/1M out

gemini-2.5-pro-preview-03-25-search

by Google

Integrated with Google's official search function.

$1.25/1M in · $10/1M out

qwen3-max-preview

by Qwen

Qwen3-Max-Preview is the latest preview model in the Qwen3 series. This version is…

$0.85/1M in · $3.38/1M out

qwen3-max

by Qwen

The Tongyi Qianwen 3 series Max model has undergone special upgrades in intelligent agent…

$0.45/1M in · $1.8/1M out
262,144 tokens context

qwen3-next-80b-a3b-instruct

by Qwen

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned model in the Qwen3-Next series…

$0.14/1M in · $0.55/1M out
256,000 tokens context

qwen3-next-80b-a3b-thinking

by Qwen

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that…

$0.14/1M in · $1.42/1M out
256,000 tokens context

qwen3-235b-a22b-instruct-2507

by Qwen

Qwen3-235B-A22B-Instruct-2507

$0.28/1M in · $1.12/1M out
262,144 tokens context

qwen3-235b-a22b-thinking-2507

by Qwen

The open-source thinking model based on Qwen3 has significantly improved in logical…

$0.28/1M in · $2.8/1M out
262,144 tokens context

qwen3-coder-30b-a3b-instruct

by Qwen

The code generation model based on Qwen3 has powerful Coding Agent capabilities…

$0.2/1M in · $0.8/1M out
2,000,000 tokens context

qwen3-coder-480b-a35b-instruct

by Qwen

The code generation model based on Qwen3 has powerful Coding Agent capabilities…

$0.82/1M in · $3.28/1M out
262,000 tokens context

DeepSeek-V3

by DeepSeek

It has been automatically upgraded to the latest released version, 250324. Automatically…

$0.27/1M in · $1.09/1M out
1,638,000 tokens context

LongCat-Flash-Chat

by Meituan

Meituan has officially released and open-sourced LongCat-Flash-Chat, which utilizes an…

$0.14/1M in · $0.7/1M out

gemini-2.5-pro-preview-06-05-search

by Google

Integrated with Google's official search function.

$1.25/1M in · $10/1M out

jina-embeddings-v5-text-nano

by Jina AI

A 3.8-billion-parameter general vector model (embedding model) for state-of-the-art…

$0.05/1M in

jina-embeddings-v5-text-small

by Jina AI

A 3.8-billion-parameter general vector (embedding) model providing state-of-the-art…

$0.05/1M in

qwen3-235b-a22b

by Qwen

Qwen3-235B-A22B is a massive 235B parameter Mixture-of-Experts (MoE) model that operates…

$0.28/1M in · $1.12/1M out
131,100 tokens context

qwen3-coder-flash

by Qwen

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3…

$0.14/1M in · $0.54/1M out
256,000 tokens context

qwen3-coder-plus

by Qwen

The code generation model based on Qwen3 has powerful Coding Agent capabilities, excels…

$0.54/1M in · $2.16/1M out
1,048,576 tokens context

qwen3-coder-plus-2025-07-22

by Qwen

The code generation model based on Qwen3 has powerful Coding Agent capabilities, excels…

$0.54/1M in · $2.16/1M out
128,000 tokens context

Qwen2.5-VL-72B-Instruct

by Qwen

The model provider is the Sophon platform. Qwen2.5-VL-72B-Instruct is the latest…

$0.62/1M in · $0.62/1M out

ernie-5.0-thinking-preview

by Baidu

The new generation Wenxin model, Wenxin 5.0, is a native full-modal large model that…

$0.82/1M in · $3.29/1M out
183,000 tokens context

inclusionAI/Ling-1T

by InclusionAI

Ling-1T is the first flagship non-thinking model in the “Ling 2.0” series, featuring 1…

$0.55/1M in · $2.19/1M out

inclusionAI/Ring-1T

by InclusionAI

Ring-1T is an open-source idea model with a trillion parameters released by the Bailing…

$0.55/1M in · $2.19/1M out

bce-reranker-base

by Qwen

Based on the dense foundational model of the Qwen3 series, it is specifically designed…

$0.07/1M in

codex-mini-latest

by OpenAI

Only supports v1/responses API calls.https://docs.aihubmix.com/en/api/Responses-API codex-…

$1.5/1M in · $6/1M out

doubao-seedream-4-0

by Doubao

Seedream 4.0 is a SOTA-level multimodal image creation model based on leading…

$2/1M in

embedding-v1

by Baidu

Embedding-V1 is a text representation model based on Baidu's Wenxin large model…

$0.07/1M in

ernie-4.5-turbo-latest

by Baidu

Wenxin 4.5 Turbo also has significant improvements in hallucination reduction, logical…

$0.11/1M in · $0.44/1M out
135,000 tokens context

glm-4.5-x

by ChatGLM

GLM-4.5-X is the high-speed version of GLM-4.5, offering powerful performance with a…

$2.2/1M in · $8.91/1M out

gme-qwen2-vl-2b-instruct

by Qwen

The GME-Qwen2VL series is a unified multimodal Embedding model trained based on the…

$0.14/1M in · $0.14/1M out

gte-rerank-v2

by Qwen

gte-rerank-v2 is a multilingual unified text ranking model developed by Tongyi Lab…

$0.11/1M in · $0.11/1M out

inclusionAI/Ling-flash-2.0

by InclusionAI

Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion…

$0.14/1M in · $0.54/1M out

inclusionAI/Ling-mini-2.0

by InclusionAI

Ling-mini-2.0 is a small-sized, high-performance large language model based on the MoE…

$0.07/1M in · $0.27/1M out

inclusionAI/Ring-flash-2.0

by InclusionAI

Ring-flash-2.0 is a high-performance thinking model deeply optimized based on the…

$0.14/1M in · $0.54/1M out

jina-deepsearch-v1

by Jina AI

DeepSearch combines search, reading, and reasoning capabilities to pursue the best…

$0.05/1M in · $0.05/1M out
1,000,000 tokens context

jina-embeddings-v4

by Jina AI

A general-purpose vector model with 3.8 billion parameters, used for multimodal and…

$0.05/1M in · $0.05/1M out

jina-reranker-v3

by Jina AI

Multimodal multilingual document reranker, 131K context, 0.6B parameters, for visual…

$0.05/1M in · $0.05/1M out
131,000 tokens context

llama-4-maverick

by Llama

Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring…

$0.2/1M in · $0.2/1M out
1,048,576 tokens context

llama-4-scout

by Llama

Llama 4 Scout is a highly efficient Mixture-of-Experts (MoE) model from Meta, activating…

$0.2/1M in · $0.2/1M out
131,000 tokens context

qwen-image

by Qwen

Qwen-Image is a foundational image generation model in the Qwen series, achieving…

$2/1M in

qwen-image-edit

by Qwen

Qwen-Image-Edit is the image editing version of Qwen-Image. Based on the 20B Qwen-Image…

$2/1M in

qwen-image-max

by Qwen

Qwen-Image-Edit is the image editing version of Qwen-Image. Based on the 20B Qwen-Image…

$2/1M in

qwen-mt-plus

by Qwen

Based on the comprehensive upgrade of Qwen3, this flagship translation large model…

$0.49/1M in · $1.48/1M out
16,000 tokens context

qwen-mt-turbo

by Qwen

Based on the comprehensive upgrade of Qwen3, this flagship translation large model…

$0.19/1M in · $0.53/1M out
16,000 tokens context

qwen3-embedding-0.6b

by Qwen

The Qwen3 Embedding model series is the latest proprietary model family from Qwen…

$0.07/1M in

qwen3-embedding-4b

by Qwen

The Qwen3 Embedding model series is the latest proprietary model family from Qwen…

$0.07/1M in · $0.07/1M out

qwen3-embedding-8b

by Qwen

The Qwen3 Embedding model series is the latest proprietary model family from Qwen…

$0.07/1M in

qwen3-reranker-0.6b

by Qwen

Based on the dense foundational model of the Qwen3 series, it is specifically designed…

$0.11/1M in · $0.11/1M out
16,000 tokens context

qwen3-reranker-4b

by Qwen

Based on the dense foundational model of the Qwen3 series, it is specifically designed…

$0.11/1M in · $0.11/1M out

qwen3-reranker-8b

by Qwen

Based on the dense foundational model of the Qwen3 series, it is specifically designed…

$0.11/1M in · $0.11/1M out

tao-8k

by Baidu

tao-8k是由Huggingface开发者amu研发并开源的长文本向量表示模型,支持8k上下文长度,模型效果在C-MTEB上居前列,是当前最优的中文长文本embeddings模型…

$0.07/1M in · $0.07/1M out

jina-clip-v2

by Jina AI

Multi-modal Embeddings Model, multilingual, 1024-dimensional, 865M parameters.

$0.05/1M in · $0.05/1M out

jina-reranker-m0

by Jina AI

Multimodal multilingual document reranker, 10K context, 2.4B parameters, for visual…

$0.05/1M in · $0.05/1M out

jina-colbert-v2

by Jina AI

Multi-language ColBERT embeddings model, 560M parameters, used for embedding and…

$0.05/1M in · $0.05/1M out

DeepSeek-R1

by DeepSeek

DeepSeek R1 is a new open-source model with performance on par with OpenAI's o1 and…

$0.4/1M in · $2/1M out
1,638,000 tokens context

gpt-4o-search-preview

by OpenAI

Using the Chat Completions API, you can directly access the fine-tuned models and tool…

$2.5/1M in · $10/1M out
128,000 tokens context

gpt-4o-mini-search-preview

by OpenAI

Using the Chat Completions API, you can directly access the fine-tuned models and tool…

$0.15/1M in · $0.6/1M out
128,000 tokens context

jina-embeddings-v3

by Jina AI

Text Embeddings Model, multilingual, 1024-dimensional, 570M parameters.

$0.05/1M in · $0.05/1M out

claude-3-7-sonnet

by Anthropic

Support for the thinking parameter through the original Claude SDK.

$3.3/1M in · $16.5/1M out
200,000 tokens context

ernie-4.5

by Baidu

Wenxin Large Model 4.5 is a next-generation native multimodal foundational model…

$0.07/1M in · $0.27/1M out
160,000 tokens context

ernie-4.5-turbo-vl

by Baidu

The new version of the Wenxin Yiyan large model significantly improves capabilities in…

$0.4/1M in · $1.2/1M out
139,000 tokens context

mimo-v2-flash-free

by Xiaomi

MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It adopts…

256,000 tokens context

FLUX-1.1-pro

by Flux

FLUX-1.1-pro is an AI image generation tool for professional creators and content…

$2/1M in · $2/1M out

o3-mini

by OpenAI

OpenAI's latest fast inference model excels at STEAM tasks and offers exceptional…

$1.1/1M in · $4.4/1M out
200,000 tokens context

doubao-seed-1-6

by Doubao

Doubao-Seed-1.6 is a brand new multimodal deep reasoning model that supports four types…

$0.18/1M in · $1.8/1M out
256,000 tokens context

doubao-seed-1-6-flash

by Doubao

Doubao-Seed-1.6-flash is an extremely fast multimodal deep thinking model, with TPOT…

$0.04/1M in · $0.44/1M out
256,000 tokens context

doubao-seed-1-6-lite

by Doubao

Doubao-Seed-1.6-lite is a brand new multimodal deep reasoning model that supports…

$0.08/1M in · $0.66/1M out
256,000 tokens context

doubao-seed-1-6-thinking

by Doubao

The Doubao-Seed-1.6-thinking model has significantly enhanced reasoning capabilities…

$0.18/1M in · $1.8/1M out
256,000 tokens context

qwen3-30b-a3b-instruct-2507

by Qwen

Significantly improved performance on reasoning tasks, including logical reasoning…

$0.1/1M in · $0.41/1M out

qwen3-30b-a3b-thinking-2507

by Qwen

Significantly improved performance on reasoning tasks, including logical reasoning…

$0.12/1M in · $1.2/1M out

Qwen2-VL-72B-Instruct

by Qwen

The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration…

$2.18/1M in · $6.54/1M out

Qwen2-VL-7B-Instruct

by Qwen

The model provider is the Sophnet platform. Qwen2-VL-7B-Instruct is the latest…

$0.28/1M in · $0.7/1M out

cc-kimi-for-coding

by Moonshot AI

for claude code

$0.2/1M in · $0.2/1M out

gemini-embedding-001

by Google

Latest version

$0.15/1M in · $0.15/1M out

gpt-oss-120b

by OpenAI

gpt-oss-120b is a 117B-parameter open-weight Mixture-of-Experts (MoE) language model from…

$0.18/1M in · $0.9/1M out
131,072 tokens context

qwen-3-235b-a22b-thinking-2507

by Qwen

cerebras

$0.28/1M in · $2.8/1M out

Qwen/Qwen3-30B-A3B

by Qwen

Provided by chutes.ai

$1/1M in · $1/1M out

Qwen/Qwen3-32B

by Qwen
$0.4/1M in · $0.8/1M out

qwen3-32b

by Qwen
$0.16/1M in · $0.64/1M out

Qwen/Qwen3-14B

by Qwen

Provided by chutes.ai

$0.5/1M in · $0.5/1M out

Qwen/Qwen3-8B

by Qwen

Provided by chutes.ai

$0.2/1M in · $0.2/1M out

embedding-2

by 智谱 ChatGLM

A text vector model that converts input text information into vector representations so…

$0.07/1M in · $0.07/1M out
8,000 tokens context

embedding-3

by 智谱 ChatGLM

A text vector model that converts input text into vector representations to work with a…

$0.07/1M in · $0.07/1M out
8,000 tokens context

gemini-2.5-pro-preview-06-05

by Google

Google’s latest multimodal flagship model, combining exceptional coding and reasoning…

$1.25/1M in · $10/1M out
1,048,576 tokens context

Qwen/Qwen2.5-VL-72B-Instruct

by Qwen

Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong…

$0.5/1M in · $0.5/1M out

o1

by OpenAI

OpenAI's most powerful O-series model supports official cache hits that halve the input…

$15/1M in · $60/1M out

o1-pro

by OpenAI

The o1 series of models are trained with reinforcement learning to think before they…

$170/1M in · $680/1M out

ByteDance-Seed/Seed-OSS-36B-Instruct

by Doubao

Seed-OSS is a series of open-source large language models developed by ByteDance's Seed…

$0.2/1M in · $0.53/1M out
256,000 tokens context

doubao-seed-1-6-250615

by Doubao

Doubao-Seed-1.6 is a brand new multimodal deep reasoning model that supports four types…

$0.18/1M in · $2.52/1M out

doubao-seed-1-6-flash-250615

by Doubao

Doubao-Seed-1.6-flash is an extremely fast multimodal deep thinking model, with TPOT…

$0.04/1M in · $0.44/1M out

doubao-seed-1-6-thinking-250615

by Doubao

The Doubao-Seed-1.6-thinking model has significantly enhanced reasoning capabilities…

$0.18/1M in · $2.52/1M out

doubao-seed-1-6-vision-250815

by Doubao

Doubao-Seed-1.6-vision is a visual deep-thinking model that demonstrates stronger general…

$0.11/1M in · $1.1/1M out

Doubao-1.5-thinking-pro

by Doubao

Doubao-1.5 is a brand-new deep thinking model that excels in specialized fields such as…

$0.62/1M in · $2.48/1M out

cc-minimax-m2

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

deepseek-ai/DeepSeek-Prover-V2-671B

by DeepSeek

Provided by chutes.ai DeepSeek Prover V2 is a 671B parameter model, speculated to be…

$0.1/1M in · $0.1/1M out

gemini-2.5-flash-preview-tts

by Google

Gemini 2.5 Flash Preview TTS is a lightweight, low-latency text-to-speech model designed…

$0.5/1M in · $0.5/1M out

gemini-2.5-pro-preview-tts

by Google

Gemini 2.5 Pro Preview TTS is a high-fidelity text-to-speech model designed for premium…

$1/1M in · $1/1M out

gemma-3-12b-it

by Google

Gemma 3 models are multimodal, handling text and image input and generating text output…

$0.2/1M in · $0.2/1M out

gemma-3-27b-it

by Google

Gemma 3 models are multimodal, handling text and image input and generating text output…

$0.2/1M in · $0.2/1M out

gemma-3-4b-it

by Google

Gemma 3 models are multimodal, handling text and image input and generating text output…

$0.2/1M in · $0.2/1M out

gemma-3n-e4b-it

by Google

Gemma 3n is a generative AI model optimized for use in everyday devices, such as phones…

$0.2/1M in · $0.2/1M out

gemma-3-1b-it

by Google

Gemma 3 models are multimodal, handling text and image input and generating text output…

$0.2/1M in · $0.2/1M out

deepseek-r1-distill-llama-70b

by DeepSeek

Provided by Groq, the DeepSeek-R1-Distill model is fine-tuned based on an open-source…

$0.8/1M in · $1.6/1M out

gpt-4o-mini-tts

by OpenAI

OpenAI’s latest TTS model, gpt-4o-mini-tts, uses the same API endpoint (/v1/audio/speech)…

$0.6/1M in · $12/1M out

tngtech/DeepSeek-R1T-Chimera

by DeepSeek

Provided by chutes.ai DeepSeek-R1T-Chimera merges DeepSeek-R1’s reasoning strengths with…

$0.02/1M in · $0.02/1M out

veo-2.0-generate-001

by Google

Veo 2.0 is an advanced video generation model capable of producing high-quality videos…

$2/1M in · $2/1M out

o1-preview

by OpenAI

The latest and most powerful inference model from OpenAI; AiHubMix uses both OpenAI and…

$15/1M in · $60/1M out

o1-mini

by OpenAI

o1-mini is faster and 80% cheaper, and is competitive with o1-preview on coding tasks…

$3/1M in · $12/1M out

gpt-4o-2024-11-20

by OpenAI

The latest version of the GPT-4o model; it is recommended to use this version, as it is…

$2.5/1M in · $10/1M out
128,000 tokens context

gpt-4o

by OpenAI

GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more…

$2.5/1M in · $10/1M out
128,000 tokens context

gpt-4o-mini

by OpenAI

The lightweight version of GPT-4o, which is affordable and fast, suitable for handling…

$0.15/1M in · $0.6/1M out
128,000 tokens context

AiHubmix-mistral-medium

by Mistral

Mistral Medium 3 is a SOTA & versatile model designed for a wide range of tasks…

$0.4/1M in · $2/1M out

ERNIE-X1.1-Preview

by Baidu

The Wenxin large model X1.1 has made significant improvements in question answering, tool…

$0.14/1M in · $0.54/1M out
119,000 tokens context

Qwen/QwQ-32B

by Qwen

Silicon-based flow provision

$0.14/1M in · $0.56/1M out

chutesai/Mistral-Small-3.1-24B-Instruct-2503

by Mistral

Mistral's latest open-source small model; provided by chutes.ai.

$0.2/1M in · $0.8/1M out

ernie-x1.1-preview

by Baidu

The Wenxin large model X1.1 has made significant improvements in question answering, tool…

$0.14/1M in · $0.54/1M out

minimax-m2

by Minimax

MiniMax-M2 redefines efficiency for intelligent agents. It is a compact, fast, and…

$0.29/1M in · $1.15/1M out
204,800 tokens context

MiniMaxAI/MiniMax-M1-80k

by Minimax

MiniMax-M1 is an open-source large-scale hybrid attention model with 456B total…

$0.6/1M in · $2.4/1M out

Qwen/Qwen2.5-VL-32B-Instruct

by Qwen

Qwen2.5-VL-32B-Instruct is an advanced multimodal model from the Tongyi Qianwen team that…

$0.24/1M in · $0.24/1M out

baidu/ERNIE-4.5-300B-A47B

by Baidu

ERNIE-4.5-300B-A47B is a large language model developed by Baidu based on a Mixture of…

$0.32/1M in · $1.28/1M out

bge-large-en

by BAAI

bge-large-en, open-sourced by the Beijing Academy of Artificial Intelligence (BAAI), is…

$0.07/1M in · $0.07/1M out

bge-large-zh

by BAAI

bge-large-zh, open-sourced by the Beijing Academy of Artificial Intelligence (BAAI), is…

$0.07/1M in · $0.07/1M out

codestral-latest

by Mistral

Mistral has launched a new code model - Codestral 25.01…

$0.4/1M in · $1.2/1M out

ernie-4.5-0.3b

by Baidu

Wenxin Large Model 4.5 is a next-generation native multimodal foundational large model…

$0.01/1M in · $0.05/1M out

ernie-4.5-turbo-128k-preview

by Baidu

Wenxin 4.5 Turbo also shows significant enhancements in reducing hallucinations, logical…

$0.11/1M in · $0.43/1M out

ernie-x1-turbo

by Baidu

Wenxin Large Model X1 possesses enhanced abilities in understanding, planning…

$0.14/1M in · $0.54/1M out
50,500 tokens context

kat-dev

by Qwen

KAT-Dev (32B) is an open-source 32B parameter model specifically designed for software…

$0.14/1M in · $0.55/1M out
128,000 tokens context

llama-3.3-70b

by Llama

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and…

$0.6/1M in · $0.6/1M out
65,536 tokens context

moonshotai/Kimi-Dev-72B

by Moonshot AI

Kimi-Dev-72B is a new generation open-source programming large model that achieved a…

$0.32/1M in · $1.28/1M out

moonshotai/Moonlight-16B-A3B-Instruct

by Moonshot AI

Provided by chutes.ai.

$0.2/1M in · $0.2/1M out

nvidia-nemotron-3-super-120b-a12b

by Nvidia

An open-source, efficient hybrid Mamba-Transformer MoE model that supports a context…

$0.11/1M in · $0.55/1M out
1,000,000 tokens context

o1-global

by OpenAI

OpenAI new model

$15/1M in · $60/1M out

qianfan-qi-vl

by Baidu

The Qianfan-QI-VL model is a proprietary image quality inspection and visual…

$0.2/1M in · $0.6/1M out

qwen2.5-vl-72b-instruct

by Qwen

Strong capability in Chinese domain recognition, comparable to ChatGPT-4.0.

$2.4/1M in · $7.2/1M out

tencent/Hunyuan-A13B-Instruct

by Hunyuan

Hunyuan-A13B-Instruct has 8 billion parameters and can match larger models by activating…

$0.14/1M in · $0.56/1M out

unsloth/gemma-3-27b-it

by Google

Google's latest open-source model; provided by chutes.ai

$0.22/1M in · $0.22/1M out

gemini-exp-1206

by Google

Google's latest experimental model, currently Google's most powerful model.

$1.25/1M in · $5/1M out

gpt-4o-zh

by OpenAI

输入任何语言自动翻译为英文给模型,模型输出内容自动翻译为中文返回;测试阶段不支持高,仅支持文本输入;

$2.5/1M in · $10/1M out

qwen-qwq-32b

by Qwen
$0.4/1M in · $0.8/1M out

unsloth/gemma-3-12b-it

by Google

Provided by chutes.ai.

$0.2/1M in · $0.8/1M out

qwen-max-0125

by Qwen

Qwen 2.5-Max latest model

$0.38/1M in · $1.52/1M out

BAAI/bge-large-en-v1.5

by BAAI

BAAI/bge-large-en-v1.5 is a large English text embedding model and part of the BGE (BAAI…

$0.03/1M in · $0.03/1M out

BAAI/bge-large-zh-v1.5

by BAAI

BAAI/bge-large-zh-v1.5 is a large Chinese text embedding model and part of the BGE (BAAI…

$0.03/1M in · $0.03/1M out

BAAI/bge-reranker-v2-m3

by BAAI

BAAI/bge-reranker-v2-m3 is a lightweight multilingual reranking model. It is developed…

$0.03/1M in · $0.03/1M out

tencent/Hunyuan-MT-7B

by Hunyuan

Hunyuan-MT-7B is a lightweight translation model with 7 billion parameters, designed to…

$0.2/1M in · $0.2/1M out

V3

by Ideogram

Fast and high-quality — top image quality in just 11 seconds per piece, with almost no…

$2/1M in · $2/1M out

V_2

by Ideogram

The Ideogram AI drawing interface is now live. This model boasts powerful text-to-image…

$2/1M in · $2/1M out

V_2_TURBO

by Ideogram

The Ideogram AI drawing interface is now live. This model boasts powerful text-to-image…

$2/1M in · $2/1M out

V_2A

by Ideogram

The Ideogram AI drawing interface is now live. This model boasts powerful text-to-image…

$2/1M in · $2/1M out

V_2A_TURBO

by Ideogram

The Ideogram AI drawing interface is now live. This model boasts powerful text-to-image…

$2/1M in · $2/1M out

V_1

by Ideogram

V_1 is a text-to-image model in the Ideogram series. It delivers strong text rendering…

$2/1M in · $2/1M out

V_1_TURBO

by Ideogram

The Ideogram AI drawing interface is now live. This model boasts powerful text-to-image…

$2/1M in · $2/1M out

doubao-embedding-large-text-240915

by Doubao

doubao-embedding-large-text-240915 Doubao Embedding is a semantic vectorization model…

$0.1/1M in · $0.1/1M out

kimi-thinking-preview

by Moonshot AI

The latest kimi model.

$30/1M in · $30/1M out

gpt-4o-2024-08-06

by OpenAI

Supports caching, with automatic halving of charges upon a cache hit.

$2.5/1M in · $10/1M out

qwen-plus-2025-07-28

by Qwen

The Tongyi Qianwen series balanced capability model has inference performance and speed…

$0.11/1M in · $1.13/1M out

qwen-plus-latest

by Qwen

The Qwen series models with balanced capabilities have inference performance and speed…

$0.11/1M in · $1.13/1M out

sonar

by Perplexity

Latest Perplexity Model

$1.6/1M in · $1.6/1M out

stepfun-ai/step3

by StepFun

Step3 is a multimodal reasoning model released by StepFun. It uses a Mixture‑of‑Experts…

$1.1/1M in · $2.75/1M out

text-embedding-v4

by Qwen

This is the Tongyi Laboratory's multilingual unified text vector model trained based on…

$0.08/1M in · $0.08/1M out

AiHubmix-Phi-4-mini-reasoning

by Microsoft

Phi-4-mini-reasoning is a lightweight open model designed for advanced mathematical…

$0.12/1M in · $0.12/1M out
128,000 tokens context

qwen-turbo-latest

by Qwen

The Qwen series model with the fastest speed and lowest cost, suitable for simple tasks…

$0.05/1M in · $0.09/1M out

aihub-Phi-4-multimodal-instruct

by Microsoft

Microsoft's latest model

$0.12/1M in · $0.48/1M out
128,000 tokens context

qwen3-30b-a3b

by Qwen

Achieves effective integration of thinking and non-thinking modes, allowing mode…

$0.12/1M in · $1.2/1M out

aihub-Phi-4-mini-instruct

by Microsoft

Microsoft's latest model

$0.12/1M in · $0.48/1M out
128,000 tokens context

grok-3

by Grok

Grok's latest model

$3/1M in · $15/1M out

aihub-Phi-4

by Microsoft

Phi-4 is a state-of-the-art open model based on a combination of synthetic datasets…

$0.12/1M in · $0.48/1M out
16,400 tokens context

claude-3-opus-20240229

by Anthropic

Claude’s previous generation strongest model

$16.5/1M in · $82.5/1M out

dall-e-3

by OpenAI

dall-e-3 is an AI image generation model that converts natural language prompts into…

$40/1M in · $40/1M out

doubao-embedding-text-240715

by Doubao

doubao-embedding-text-240715 Doubao Embedding is a semantic vectorization model developed…

$0.7/1M in · $0.7/1M out

grok-3-beta

by Grok

Grok's latest model This model ID with beta has been officially taken offline. Using this…

$3/1M in · $15/1M out

qwen3-14b

by Qwen

Achieves effective integration of thinking and non-thinking modes, enabling mode…

$0.16/1M in · $1.6/1M out

grok-3-fast

by Grok
$5.5/1M in · $27.5/1M out

qwen3-8b

by Qwen

Achieves effective integration of thinking and non-thinking modes, enabling mode…

$0.08/1M in · $0.8/1M out

deepseek-ai/DeepSeek-R1-Zero

by DeepSeek

Openly deployed by chutes.ai; inference with FP8; zero is the initial preliminary version…

$2.2/1M in · $2.2/1M out

grok-3-fast-beta

by Grok
$5.5/1M in · $27.5/1M out

grok-3-mini

by Grok
$0.3/1M in · $0.5/1M out

qwen3-4b

by Qwen

Achieves effective integration of thinking and non-thinking modes, allowing mode…

$0.05/1M in · $0.46/1M out

grok-3-mini-beta

by Grok

This model ID with beta has been officially taken offline. Using this model…

$0.33/1M in · $0.55/1M out

qwen3-1.7b

by Qwen

Effectively integrates thinking and non-thinking modes, allowing mode switching during…

$0.05/1M in · $0.46/1M out

qwen3-0.6b

by Qwen

Effectively integrates thinking and non-thinking modes, allowing mode switching during…

$0.05/1M in · $0.46/1M out

alicloud-glm-5

by Z.AI

GLM-5 is an advanced, open-source large language model designed for developers tackling…

$0.56/1M in · $2.54/1M out

command-a-03-2025

by Cohere

Command A is Cohere most performant model to date, excelling at tool use, agents…

$2.5/1M in · $10/1M out

grok-3-mini-fast-beta

by Grok
$0.33/1M in · $2.2/1M out

qwen-3-32b

by Qwen

cerebras

$0.4/1M in · $1.6/1M out

qwen-turbo-2025-04-28

by Qwen

The Qwen3 series Turbo model effectively integrates thinking and non-thinking modes…

$0.05/1M in · $0.09/1M out

qwen-plus-2025-04-28

by Qwen

The Qwen3 series Plus model effectively integrates thinking and non-thinking modes…

$0.11/1M in · $1.13/1M out

THUDM/GLM-Z1-32B-0414

by ChatGLM

GLM-Z1-32B-0414 is a reasoning-focused AI model built on GLM-4-32B-0414. It has been…

$0.08/1M in · $0.08/1M out

THUDM/GLM-4.1V-9B-Thinking

by ChatGLM

GLM-4.1V-9B-Thinking is an open-source Vision Language Model (VLM) jointly released by…

$0.1/1M in · $0.1/1M out

text-embedding-004

by Google
$0.02/1M in · $0.02/1M out

THUDM/GLM-4-32B-0414

by ChatGLM

GLM-4-32B-0414 is a next-generation open-source model with 32 billion parameters…

$0.08/1M in · $0.08/1M out

THUDM/GLM-Z1-9B-0414

by ChatGLM

GLM-Z1-9B-0414 is a small but powerful model in the GLM series, with only 9 billion…

$0.05/1M in · $0.05/1M out

THUDM/GLM-4-9B-0414

by ChatGLM

GLM-4-9B-0414 is a lightweight model in the GLM family, with 9 billion parameters. It…

$0.05/1M in · $0.05/1M out

cc-doubao-seed-code-preview-latest

by Doubao

claude code

$0.2/1M in · $0.2/1M out

doubao-seed-code-preview-latest

by Doubao

chat

$0.2/1M in · $0.2/1M out

deepseek-ai/Janus-Pro-7B

by DeepSeek

Janus-Pro deepseek最新发布图片生成模型;是一个新颖的自回归框架,统一了多模态理解和生成。它通过将视觉编码解耦为独立路径来解决以往方法的局限性,同时仍然利用单一的统…

$2/1M in · $2/1M out

glm-zero-preview

by ChatGLM

Simply put, it is the intelligent enhanced version of O1.

$2/1M in · $2/1M out

qwen-3-235b-a22b-instruct-2507

by Qwen

cerebras

$0.28/1M in · $1.4/1M out

coding-glm-4.5-air

by ChatGLM
$0.01/1M in · $0.08/1M out

deepinfra-nvidia-nemotron-3-nano-30b-a3b2

by Nvidia
$0.07/1M in · $0.26/1M out

glm-4.5-air

by ChatGLM
$0.14/1M in · $0.84/1M out
131,072 tokens context

gpt-4-32k

by OpenAI

The smartest version of GPT-4; OpenAI no longer offers it officially. All the 32k…

$60/1M in · $120/1M out

nvidia-llama-3.1-nemotron-70b-instruct

by Nvidia
$1.32/1M in · $1.32/1M out

nvidia-llama-3.3-nemotron-super-49b-v1.5

by Nvidia
$0.11/1M in · $0.44/1M out

nvidia-nemotron-3-nano-30b-a3b

by Nvidia
$0.07/1M in · $0.26/1M out

nvidia-nemotron-nano-12b-v2-vl

by Nvidia
$0.22/1M in · $0.66/1M out

nvidia-nemotron-nano-9b-v2

by Nvidia
$0.04/1M in · $0.18/1M out

o1-preview-2024-09-12

by OpenAI
$15/1M in · $60/1M out

Qwen/QVQ-72B-Preview

by Qwen
$1.2/1M in · $1.2/1M out

Qwen/QwQ-32B-Preview

by Qwen
$0.16/1M in · $0.16/1M out

llama-3.1-sonar-huge-128k-online

by Perplexity

On February 22, 2025, this model will be officially discontinued. The Perplexity AI…

$5.6/1M in · $5.6/1M out

aihubmix-Mistral-Large-2411

by Mistral

The latest Mistral Large 2 model is deployed on Azure.

$2/1M in · $6/1M out

llama-3.1-sonar-large-128k-online

by Perplexity

On February 22, 2025, this model will be officially discontinued; Perplexity AI's…

$1.2/1M in · $1.2/1M out

aihubmix-Mistral-large-2407

by Mistral
$3/1M in · $9/1M out

grok-2-1212

by Grok

马斯克的xai的最新模型,aihubmix价格低于官网10%

$1.8/1M in · $9/1M out

llama-3.1-70b

by Llama
$0.44/1M in · $0.44/1M out

wan2.6-t2i

by Qwen
$2/1M in

DESCRIBE

by Ideogram

This endpoint is used to describe an image. Supported image formats include JPEG, PNG…

$2/1M in · $2/1M out

UPSCALE

by Ideogram

The super-resolution upscale interface of the Ideogram AI drawing model is designed to…

$2/1M in · $2/1M out

bai-qwen3-vl-235b-a22b-instruct

by Qwen

The Qwen3 series open-source models include hybrid models, thinking models, and…

$0.27/1M in · $1.1/1M out

cc-MiniMax-M2

by Minimax

For Claude Code only

$0.1/1M in · $0.1/1M out

cc-deepseek-v3

by DeepSeek

For Claude code only

$0.3/1M in · $0.3/1M out

cc-deepseek-v3.1

by DeepSeek

For Claude code only

$0.56/1M in · $1.68/1M out

cc-ernie-4.5-300b-a47b

by Baidu

For Claude code only

$0.32/1M in · $1.28/1M out

cc-kimi-dev-72b

by Moonshot AI

For Claude code only

$0.32/1M in · $1.28/1M out

cc-kimi-k2-instruct

by Moonshot AI

For Claude code only

$1.1/1M in · $3.3/1M out

cc-kimi-k2-instruct-0905

by Moonshot AI

For Claude code only

$1.1/1M in · $3.3/1M out

cc-kimi-k2-thinking

by Moonshot AI

Dedicated for Claude Code

$0.55/1M in · $2.19/1M out

computer-use-preview

by OpenAI
$3/1M in · $12/1M out

gpt-image-test

by OpenAI
$5/1M in · $40/1M out

grok-4.20-beta-0309-non-reasoning

by Grok

Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent…

$2/1M in · $6/1M out
2,000,000 tokens context

grok-4.20-beta-0309-reasoning

by Grok

Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent…

$2/1M in · $6/1M out
2,000,000 tokens context

grok-4.20-multi-agent-beta-0309

by Grok

Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent…

$2/1M in · $6/1M out
2,000,000 tokens context

jina-reader

by Jina AI
$0.05/1M in · $0.05/1M out

jina-search

by Jina AI
$0.05/1M in · $0.05/1M out

llama3.1-8b

by Llama

cerebras

$0.3/1M in · $0.6/1M out

o1-2024-12-17

by OpenAI
$15/1M in · $60/1M out

sf-kimi-k2-thinking

by Moonshot AI
$0.55/1M in · $2.19/1M out

Baichuan3-Turbo

by Baichuan
$1.9/1M in · $1.9/1M out

Baichuan3-Turbo-128k

by Baichuan
$3.8/1M in · $3.8/1M out

Baichuan4

by Baichuan
$16/1M in · $16/1M out

Baichuan4-Air

by Baichuan
$0.16/1M in · $0.16/1M out

Baichuan4-Turbo

by Baichuan
$2.4/1M in · $2.4/1M out

DeepSeek-v3

by DeepSeek
$0.27/1M in · $1.09/1M out

Doubao-1.5-lite-32k

by Doubao

Doubao-1.5-lite, a brand-new generation of lightweight model, offers exceptional response…

$0.05/1M in · $0.1/1M out

Doubao-1.5-pro-256k

by Doubao

Doubao-1.5-pro-256k, a fully upgraded version based on Doubao-1.5-Pro, delivers an…

$0.8/1M in · $1.44/1M out

Doubao-1.5-pro-32k

by Doubao

Doubao-1.5-pro, a brand-new generation of flagship model, features comprehensive…

$0.13/1M in · $0.34/1M out

Doubao-1.5-vision-pro-32k

by Doubao

Doubao-1.5-vision-pro is a newly upgraded multimodal large model that supports image…

$0.46/1M in · $1.38/1M out

Doubao-lite-128k

by 字节跳动
$0.14/1M in · $0.28/1M out

Doubao-lite-32k

by 字节跳动
$0.06/1M in · $0.12/1M out

Doubao-lite-4k

by 字节跳动
$0.06/1M in · $0.12/1M out

Doubao-pro-128k

by 字节跳动
$0.8/1M in · $1.44/1M out

Doubao-pro-256k

by 字节跳动
$0.8/1M in · $1.44/1M out

Doubao-pro-32k

by 字节跳动
$0.14/1M in · $0.35/1M out

Doubao-pro-4k

by 字节跳动
$0.14/1M in · $0.35/1M out

GPT-OSS-20B

by OpenAI
$0.11/1M in · $0.55/1M out

Gryphe/MythoMax-L2-13b

by Meta
$0.4/1M in · $0.4/1M out

MiniMax-Text-01

by Minimax
$0.14/1M in · $1.12/1M out

Mistral-large-2407

by Mistral
$3/1M in · $9/1M out

Qwen/Qwen2-1.5B-Instruct

by Qwen
$0.2/1M in · $0.2/1M out

Qwen/Qwen2-57B-A14B-Instruct

by Qwen
$0.24/1M in · $0.24/1M out

Qwen/Qwen2-72B-Instruct

by Qwen
$0.8/1M in · $0.8/1M out

Qwen/Qwen2-7B-Instruct

by Qwen
$0.08/1M in · $0.08/1M out

Qwen/Qwen2.5-32B-Instruct

by Qwen
$0.6/1M in · $0.6/1M out

Qwen/Qwen2.5-72B-Instruct

by Qwen
$0.8/1M in · $0.8/1M out

Qwen/Qwen2.5-72B-Instruct-128K

by Qwen
$0.8/1M in · $0.8/1M out

Qwen/Qwen2.5-7B-Instruct

by Qwen
$0.4/1M in · $0.4/1M out

Qwen/Qwen2.5-Coder-32B-Instruct

by Qwen
$0.16/1M in · $0.16/1M out

Qwen3-235B-A22B-Thinking-2507

by Qwen

算能提供

$0.28/1M in · $2.8/1M out

Stable-Diffusion-3-5-Large

by Stable diffusion

Stable Diffusion 3.5 Large, developed by Stability AI, is a text-to-image generation…

$4/1M in · $4/1M out

WizardLM/WizardCoder-Python-34B-V1.0

by Meta
$0.9/1M in · $0.9/1M out

ahm-Phi-3-5-MoE-instruct

by Microsoft

Phi-3.5-MoE 是一个轻量级的最先进开放模型,基于用于 Phi-3 的数据集构建——合成数据和经过筛选的公开可用文档,重点关注高质量、推理密集的数据。该模型支持多语言,并具…

$0.4/1M in · $1.6/1M out

ahm-Phi-3-5-mini-instruct

by Microsoft

Phi-3.5-mini is a lightweight, state-of-the-art open model built upon the dataset used…

$1/1M in · $3/1M out

ahm-Phi-3-5-vision-instruct

by Microsoft
$0.4/1M in · $1.6/1M out

ahm-Phi-3-medium-128k

by Microsoft
$6/1M in · $18/1M out

ahm-Phi-3-medium-4k

by Microsoft
$1/1M in · $3/1M out

ahm-Phi-3-small-128k

by Microsoft
$1/1M in · $3/1M out

aihubmix-Codestral-2501

by Mistral

azure部署

$0.4/1M in · $1.2/1M out

aihubmix-Cohere-command-r

by Cohere
$0.64/1M in · $1.92/1M out

aihubmix-Jamba-1-5-Large

by AI21
$2.2/1M in · $8.8/1M out

aihubmix-Llama-3-1-405B-Instruct

by Meta
$5/1M in · $15/1M out

aihubmix-Llama-3-1-70B-Instruct

by Meta
$0.6/1M in · $0.78/1M out

aihubmix-Llama-3-1-8B-Instruct

by Meta
$0.3/1M in · $0.6/1M out

aihubmix-Llama-3-2-11B-Vision

by Meta
$0.4/1M in · $0.4/1M out

aihubmix-Llama-3-2-90B-Vision

by Meta
$2.4/1M in · $2.4/1M out

aihubmix-Llama-3-70B-Instruct

by Meta
$0.7/1M in · $0.7/1M out

aihubmix-Mistral-large

by Mistral
$4/1M in · $12/1M out

aihubmix-command-r-08-2024

by Cohere
$0.2/1M in · $0.8/1M out

aihubmix-command-r-plus

by Cohere
$3.84/1M in · $19.2/1M out

aihubmix-command-r-plus-08-2024

by Cohere
$2.8/1M in · $11.2/1M out

alicloud-deepseek-v3.2

by DeepSeek
$0.27/1M in · $0.41/1M out

alicloud-glm-4.7

by Z.AI
$0.41/1M in · $1.92/1M out

alicloud-kimi-k2-thinking

by Moonshot AI
$0.55/1M in · $2.19/1M out

alicloud-kimi-k2.5

by Moonshot AI
$0.55/1M in · $2.88/1M out
256,000 tokens context

alicloud-minimax-m2.5

by Minimax
$0.29/1M in · $1.15/1M out

anthropic-opus-4-6

by Anthropic

Claude Opus 4.6 is Anthropic’s latest state-of-the-art reasoning model. It features an…

$5/1M in · $25/1M out
200,000 tokens context

azure-deepseek-v3.2

by DeepSeek
$0.58/1M in · $1.68/1M out

azure-deepseek-v3.2-speciale

by DeepSeek
$0.58/1M in · $1.68/1M out

azure-kimi-k2.5

by Moonshot AI
$0.6/1M in · $3/1M out
256,000 tokens context

cbs-glm-4.7

by Z.AI
$2.25/1M in · $2.75/1M out

cerebras-llama-3.3-70b

by Llama
$0.6/1M in · $0.6/1M out

chatglm_lite

by 智谱 ChatGLM
$0.29/1M in · $0.29/1M out

chatglm_pro

by 智谱 ChatGLM
$1.43/1M in · $1.43/1M out

chatglm_std

by 智谱 ChatGLM
$0.71/1M in · $0.71/1M out

chatglm_turbo

by 智谱 ChatGLM
$0.71/1M in · $0.71/1M out

claude-2

by Anthropic
$8.8/1M in · $8.8/1M out

claude-2.0

by Anthropic
$8.8/1M in · $39.6/1M out

claude-2.1

by Anthropic
$8.8/1M in · $39.6/1M out

claude-3-haiku-20240229

by Anthropic
$0.28/1M in · $0.28/1M out

claude-3-haiku-20240307

by Anthropic
$0.28/1M in · $1.38/1M out

claude-3-sonnet-20240229

by Anthropic
$3.3/1M in · $16.5/1M out

claude-instant-1

by Anthropic
$1.79/1M in · $1.79/1M out

claude-instant-1.2

by Anthropic
$0.88/1M in · $3.96/1M out

code-davinci-edit-001

by 智谱 ChatGLM
$20/1M in · $20/1M out

cogview-3

by 智谱 ChatGLM
$35.5/1M in · $35.5/1M out

cogview-3-plus

by 智谱 ChatGLM
$10/1M in · $10/1M out

command

by Cohere
$1/1M in · $2/1M out

command-light

by Cohere
$1/1M in · $2/1M out

command-light-nightly

by Cohere
$1/1M in · $2/1M out

command-nightly

by Cohere
$1/1M in · $2/1M out

command-r

by Cohere
$0.64/1M in · $1.92/1M out

command-r-08-2024

by Cohere
$0.2/1M in · $0.8/1M out

command-r-plus

by Cohere
$3.84/1M in · $19.2/1M out

command-r-plus-08-2024

by Cohere
$2.8/1M in · $11.2/1M out

dall-e-2

by OpenAI
$16/1M in · $16/1M out

davinci

by OpenAI
$20/1M in · $20/1M out

davinci-002

by OpenAI
$2/1M in · $2/1M out

deepinfra-llama-3.1-8b-instant

by Llama
$0.03/1M in · $0.05/1M out

deepinfra-llama-3.3-70b-instant-turbo

by Llama
$0.11/1M in · $0.35/1M out

deepinfra-llama-4-maverick-17b-128e-instruct

by Llama
$0.33/1M in · $1.32/1M out

deepinfra-llama-4-scout-17b-16e-instruct

by Llama
$0.09/1M in · $0.33/1M out

deepseek-ai/DeepSeek-Coder-V2-Instruct

by DeepSeek
$0.16/1M in · $0.32/1M out

deepseek-ai/DeepSeek-R1-Distill-Llama-70B

by DeepSeek

来自siliconflow开源部署,模型本身通过知识蒸馏得到的模型

$0.6/1M in · $0.6/1M out

deepseek-ai/DeepSeek-R1-Distill-Llama-8B

by DeepSeek

来自siliconflow开源部署,模型本身通过知识蒸馏得到的模型

$0.01/1M in · $0.01/1M out

deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

by DeepSeek

来自siliconflow开源部署,模型本身通过知识蒸馏得到的模型

$0.01/1M in · $0.01/1M out

deepseek-ai/DeepSeek-R1-Distill-Qwen-14B

by DeepSeek

Open source deployment from SiliconFlow, the model itself is obtained through knowledge…

$0.1/1M in · $0.1/1M out

deepseek-ai/DeepSeek-R1-Distill-Qwen-32B

by DeepSeek

Open source deployment from SiliconFlow, the model itself is obtained through knowledge…

$0.2/1M in · $0.2/1M out

deepseek-ai/DeepSeek-R1-Distill-Qwen-7B

by DeepSeek

Open source deployment from SiliconFlow, the model itself is obtained through knowledge…

$0.01/1M in · $0.01/1M out

deepseek-ai/DeepSeek-V2-Chat

by DeepSeek
$0.16/1M in · $0.32/1M out

deepseek-ai/DeepSeek-V2.5

by DeepSeek
$0.16/1M in · $0.32/1M out

deepseek-ai/deepseek-llm-67b-chat

by DeepSeek
$0.16/1M in · $0.16/1M out

deepseek-ai/deepseek-vl2

by DeepSeek
$0.16/1M in · $0.16/1M out

deepseek-v3

by DeepSeek
$0.27/1M in · $1.09/1M out

distil-whisper-large-v3-en

by OpenAI

Groq开源部署非

$5.56/1M in · $5.56/1M out

doubao-1-5-thinking-vision-pro-250428

by Doubao

Deep Thinking Image Understanding Visual Localization Video Understanding Tool…

$2/1M in · $2/1M out

fx-flux-2-pro

by Flux
$2/1M in

gemini-2.5-pro-exp-03-25

by Google

Google’s latest experimental model, highly unstable, for experience only. It boasts…

$1.25/1M in · $5/1M out

gemini-embedding-exp-03-07

by Google
$0.02/1M in · $0.02/1M out

gemini-exp-1114

by Google
$1.25/1M in · $5/1M out

gemini-exp-1121

by Google
$1.25/1M in · $5/1M out

gemini-pro

by Google

gemini最初版不推荐

$0.2/1M in · $0.6/1M out

gemini-pro-vision

by Google

gemini最初版不推荐

$1/1M in · $1/1M out

gemma-7b-it

by Google
$0.1/1M in · $0.1/1M out

glm-3-turbo

by 智谱 ChatGLM
$0.71/1M in · $0.71/1M out

glm-4

by 智谱 ChatGLM
$14.2/1M in · $14.2/1M out

glm-4-flash

by 智谱 ChatGLM
$0.1/1M in · $0.1/1M out

glm-4-plus

by 智谱 ChatGLM
$8/1M in · $8/1M out

glm-4.5-airx

by ChatGLM

GLM-4.5-AirX is the high-speed version of GLM-4.5-Air, with faster response times…

$1.1/1M in · $4.51/1M out

glm-4v

by 智谱 ChatGLM
$14.2/1M in · $14.2/1M out

glm-4v-plus

by 智谱 ChatGLM
$2/1M in · $2/1M out

google-gemma-3-12b-it

by Google
$0.2/1M in · $0.2/1M out

google-gemma-3-27b-it

by Google
$0.2/1M in · $0.2/1M out

google-gemma-3-4b-it

by Google
$0.2/1M in · $0.2/1M out

google/gemini-exp-1114

by Google
$1.25/1M in · $5/1M out

google/gemma-2-27b-it

by Google
$0.8/1M in · $0.8/1M out

google/gemma-2-9b-it:free

by Google
$0.02/1M in · $0.02/1M out

gpt-3.5-turbo

by OpenAI

Since the GPT-3.5-turbo model has been officially deprecated, all requests targeting this…

$0.5/1M in · $1.5/1M out

gpt-3.5-turbo-0301

by OpenAI
$1.5/1M in · $1.5/1M out

gpt-3.5-turbo-0613

by OpenAI
$1.5/1M in · $2/1M out

gpt-3.5-turbo-1106

by OpenAI
$1/1M in · $2/1M out

gpt-3.5-turbo-16k

by OpenAI
$3/1M in · $4/1M out

gpt-3.5-turbo-16k-0613

by OpenAI
$3/1M in · $4/1M out

gpt-3.5-turbo-instruct

by OpenAI
$1.5/1M in · $2/1M out

gpt-4

by OpenAI
$30/1M in · $60/1M out

gpt-4-0125-preview

by OpenAI
$10/1M in · $30/1M out

gpt-4-0314

by OpenAI
$30/1M in · $60/1M out

gpt-4-0613

by OpenAI
$30/1M in · $60/1M out

gpt-4-1106-preview

by OpenAI
$10/1M in · $30/1M out

gpt-4-32k-0314

by OpenAI
$60/1M in · $120/1M out

gpt-4-32k-0613

by OpenAI
$60/1M in · $120/1M out

gpt-4-turbo

by OpenAI
$10/1M in · $30/1M out

gpt-4-turbo-2024-04-09

by OpenAI
$10/1M in · $30/1M out

gpt-4-turbo-preview

by OpenAI
$10/1M in · $30/1M out

gpt-4-vision-preview

by OpenAI
$10/1M in · $30/1M out

gpt-4o-2024-05-13

by OpenAI
$5/1M in · $15/1M out
128,000 tokens context

gpt-4o-mini-2024-07-18

by OpenAI
$0.15/1M in · $0.6/1M out

gpt-oss-20b

by OpenAI

gpt-oss-20b is a 21-billion parameter open-weight model released by OpenAI under the…

$0.11/1M in · $0.55/1M out
128,000 tokens context

grok-2-vision-1212

by Grok

grok-2-vision-1212 is the latest vision model in the Grok family, delivering outstanding…

$1.8/1M in · $9/1M out

grok-vision-beta

by Grok
$5.6/1M in · $16.8/1M out

groq-llama-3.1-8b-instant

by Llama
$0.06/1M in · $0.09/1M out

groq-llama-3.3-70b-versatile

by Llama
$0.65/1M in · $0.87/1M out

groq-llama-4-maverick-17b-128e-instruct

by Llama
$0.22/1M in · $0.66/1M out

groq-llama-4-scout-17b-16e-instruct

by Llama
$0.12/1M in · $0.37/1M out

jina-embeddings-v2-base-code

by Jina AI

Model optimized for code and document search, 768-dimensional, 137M parameters.

$0.05/1M in · $0.05/1M out

learnlm-1.5-pro-experimental

by Google
$1.25/1M in · $5/1M out

llama-3.1-405b-instruct

by Meta
$4/1M in · $4/1M out

llama-3.1-405b-reasoning

by Meta
$4/1M in · $4/1M out

llama-3.1-70b-versatile

by Meta
$0.6/1M in · $0.6/1M out

llama-3.1-8b-instant

by Llama
$0.3/1M in · $0.6/1M out

llama-3.1-sonar-small-128k-online

by Perplexity

On February 22, 2025, this model will be officially discontinued. The Perplexity AI…

$0.3/1M in · $0.3/1M out

llama-3.2-11b-vision-preview

by Meta
$0.2/1M in · $0.2/1M out

llama-3.2-1b-preview

by Meta
$0.2/1M in · $0.2/1M out

llama-3.2-3b-preview

by Meta
$0.2/1M in · $0.2/1M out

llama-3.2-90b-vision-preview

by Meta
$2.4/1M in · $2.4/1M out

llama2-70b-4096

by Llama
$0.5/1M in · $0.5/1M out

llama2-70b-40960

by Llama
$0.5/1M in · $0.5/1M out

llama2-7b-2048

by Meta
$0.1/1M in · $0.1/1M out

llama3-70b-8192

by Meta
$0.7/1M in · $0.94/1M out

llama3-8b-8192

by Llama
$0.06/1M in · $0.12/1M out

llama3-groq-70b-8192-tool-use-preview

by Meta
$0.00089/1M in · $0.00089/1M out

llama3-groq-8b-8192-tool-use-preview

by Meta
$0.00019/1M in · $0.00019/1M out

mai-image-2

by Microsoft
$2/1M in · $2/1M out

meta-llama/Llama-3.2-90B-Vision-Instruct

by Meta
$0.5/1M in · $0.5/1M out

meta-llama/llama-3.1-405b-instruct:free

by Meta
$0.02/1M in · $0.02/1M out

meta-llama/llama-3.1-70b-instruct:free

by Meta
$0.02/1M in · $0.02/1M out

meta-llama/llama-3.1-8b-instruct:free

by Meta
$0.02/1M in · $0.02/1M out

meta-llama/llama-3.2-11b-vision-instruct:free

by Meta
$0.02/1M in · $0.02/1M out

meta-llama/llama-3.2-3b-instruct:free

by Meta
$0.02/1M in · $0.02/1M out

meta/llama-3.1-405b-instruct

by Meta
$5/1M in · $5/1M out

meta/llama3-8B-chat

by Meta
$0.3/1M in · $0.3/1M out

mistralai/mistral-7b-instruct:free

by Mistral
$0.002/1M in · $0.002/1M out

mm-minimax-m3

by Minimax
$0.29/1M in · $1.15/1M out

moonshot-kimi-k2.5

by Moonshot AI
$0.6/1M in · $3/1M out

moonshot-v1-128k

by Moonshot AI
$10/1M in · $10/1M out

moonshot-v1-128k-vision-preview

by Moonshot AI
$10/1M in · $10/1M out

moonshot-v1-32k

by Moonshot AI
$4/1M in · $4/1M out

moonshot-v1-32k-vision-preview

by Moonshot AI
$4/1M in · $4/1M out

moonshot-v1-8k

by Moonshot AI
$2/1M in · $2/1M out

moonshot-v1-8k-vision-preview

by Moonshot AI
$2/1M in · $2/1M out

nvidia/Llama-3_1-Nemotron-Ultra-253B-v1

by Nvidia

Llama-3.1-Nemotron-Ultra-253B is a 253 billion parameter reasoning-focused language model…

$0.5/1M in · $0.5/1M out

o1-mini-2024-09-12

by OpenAI
$3/1M in · $12/1M out

omni-moderation-latest

by OpenAI
$0.02/1M in · $0.02/1M out

qwen-flash

by Qwen

The model adopts tiered pricing.

$0.02/1M in · $0.2/1M out

qwen-flash-2025-07-28

by Qwen

The model adopts tiered pricing.

$0.02/1M in · $0.2/1M out

qwen-long

by Qwen
$0.1/1M in · $0.4/1M out

qwen-max

by Qwen
$0.38/1M in · $1.52/1M out

qwen-max-longcontext

by Qwen
$7/1M in · $21/1M out

qwen-plus

by Qwen
$0.11/1M in · $1.13/1M out

qwen-turbo

by Qwen
$0.05/1M in · $0.09/1M out

qwen-turbo-2024-11-01

by Qwen
$0.05/1M in · $0.09/1M out

qwen2.5-14b-instruct

by Qwen
$0.4/1M in · $1.2/1M out

qwen2.5-32b-instruct

by Qwen
$0.6/1M in · $1.2/1M out

qwen2.5-3b-instruct

by Qwen
$0.4/1M in · $0.8/1M out

qwen2.5-72b-instruct

by Qwen
$0.8/1M in · $2.4/1M out

qwen2.5-7b-instruct

by Qwen
$0.4/1M in · $0.8/1M out

qwen2.5-coder-1.5b-instruct

by Qwen
$0.2/1M in · $0.4/1M out

qwen2.5-coder-7b-instruct

by Qwen
$0.2/1M in · $0.4/1M out

qwen2.5-math-1.5b-instruct

by Qwen
$0.2/1M in · $0.2/1M out

qwen2.5-math-72b-instruct

by Qwen
$0.8/1M in · $2.4/1M out

qwen2.5-math-7b-instruct

by Qwen
$0.2/1M in · $0.4/1M out

step-2-16k

by 阶跃星辰
$2/1M in · $2/1M out

text-ada-001

by OpenAI
$0.4/1M in · $0.4/1M out

text-babbage-001

by OpenAI
$0.5/1M in · $0.5/1M out

text-curie-001

by OpenAI
$2/1M in · $2/1M out

text-davinci-002

by OpenAI
$20/1M in · $20/1M out

text-davinci-003

by OpenAI
$20/1M in · $20/1M out

text-davinci-edit-001

by OpenAI
$20/1M in · $20/1M out

text-embedding-3-large

by OpenAI
$0.13/1M in · $0.13/1M out

text-embedding-3-small

by OpenAI
$0.02/1M in · $0.02/1M out

text-embedding-ada-002

by OpenAI
$0.1/1M in · $0.1/1M out

text-embedding-v1

by OpenAI
$0.1/1M in · $0.1/1M out

text-moderation-007

by OpenAI
$0.2/1M in · $0.2/1M out

text-moderation-latest

by OpenAI
$0.2/1M in · $0.2/1M out

text-moderation-stable

by OpenAI
$0.2/1M in · $0.2/1M out

text-search-ada-doc-001

by OpenAI
$20/1M in · $20/1M out

tts-1

by OpenAI
$15/1M in · $15/1M out

tts-1-1106

by OpenAI
$15/1M in · $15/1M out

tts-1-hd

by OpenAI
$30/1M in · $30/1M out

tts-1-hd-1106

by OpenAI
$30/1M in · $30/1M out

whisper-1

by OpenAI

Ignore the displayed price on the page; the actual charge for this model request is…

$100/1M in · $100/1M out

whisper-large-v3

by OpenAI

Groq开源部署

$30.83/1M in · $30.83/1M out

whisper-large-v3-turbo

by OpenAI

Groq开源部署

$5.56/1M in · $5.56/1M out

yi-large

by 零一万物
$3/1M in · $3/1M out

yi-large-rag

by 零一万物
$4/1M in · $4/1M out

yi-large-turbo

by 零一万物
$1.8/1M in · $1.8/1M out

yi-lightning

by 零一万物
$0.2/1M in · $0.2/1M out

yi-medium

by 零一万物
$0.4/1M in · $0.4/1M out

yi-vl-plus

by 零一万物
$0.00085/1M in · $0.00085/1M out

deepseek-r1-distill-qianfan-llama-8b

by DeepSeek
$0.14/1M in · $0.55/1M out

doubao-1-5-pro-256k-250115

by 字节跳动豆包
$0.68/1M in · $1.23/1M out

doubao-1-5-pro-32k-250115

by 字节跳动豆包
$0.11/1M in · $0.27/1M out

gpt-4o-2024-08-06-global

by OpenAI
$2.5/1M in · $10/1M out

gpt-4o-mini-global

by OpenAI
$0.15/1M in · $0.6/1M out

meta-llama-3-70b

by Meta
$4.79/1M in · $4.79/1M out

meta-llama-3-8b

by Meta
$0.55/1M in · $0.55/1M out

o3-global

by OpenAI
$2/1M in · $8/1M out

o3-mini-global

by OpenAI
$1.1/1M in · $4.4/1M out

o3-pro-global

by OpenAI
$20/1M in · $80/1M out

qianfan-chinese-llama-2-13b

by Baidu
$0.82/1M in · $0.82/1M out

qianfan-llama-vl-8b

by Baidu
$0.27/1M in · $0.69/1M out