{"data":[{"model_id":"auto","model_name":"Auto","developer_id":12,"desc":"AIHubMix Smart Router: Fill in the model name as auto, and the gateway will automatically select the optimal model based on the request content;\n\nThe auto without a suffix uses the default strategy cost_optimized. You can explicitly specify the focus using auto:\u003cstrategy\u003e, such as auto (= auto:cost_optimized) cost priority: choose the cheapest as long as the capability meets the standard; auto:balanced balanced: considers capability / cost / latency\n\nauto:quality_first quality priority: prioritize selecting the strongest capability for complex reasoning and critical outputs\n\nauto:latency_critical low latency priority: prioritize selecting the fastest response","pricing":{"input":2,"output":2},"types":"llm","features":"thinking,web,deepsearch,tools,function_calling,structured_outputs,long_context","input_modalities":"text,audio,image,video","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5.3","model_name":"GLM 5.3","developer_id":5,"desc":"GLM-5.3 is Z.AI’s coding and agentic reasoning model, built for complex software engineering, long-running agent tasks, vulnerability analysis, and other demanding workloads. Building on GLM-5.2, it incorporates further post-training improvements to deliver stronger coding performance, better task execution, and greater token efficiency.\n\nWe currently offer the production-ready GLM-5.3 API with unlimited concurrency, making it well suited for high-throughput workloads, coding agents, and large-scale automation. For a limited time, GLM-5.3 is available at 10% off.","pricing":{"cache_read":0.2817,"input":1.1268,"output":3.9438},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false,"promotion":{"name":"GLM-5.3 limited-time special offer","off_percent":10,"time_type":"daily","daily":{"ranges":["00:00-23:59"]}}},{"model_id":"coding-glm-5.3","model_name":"Coding GLM 5.3","developer_id":5,"desc":"GLM-5.3 is Z.ai’s reasoning model for coding and agentic workflows, designed for complex software engineering, long-running agents, and vulnerability analysis. It uses the same base model as GLM-5.2, with scaled post-training improving coding, task execution, and token efficiency.\n\nThis model is a limited-time preview version of GLM-5.3, intended for testing and evaluation only. Service stability is not guaranteed, and we do not recommend using it in production environments. We’re waiting for the official commercial API release and will integrate it as soon as official support becomes available.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.7-flash","model_name":"Gemini 3.7 Flash","developer_id":8,"desc":"Gemini 3.7 Flash is Google’s natively multimodal reasoning model for coding, agents, web development, and knowledge work. It supports a 1M-token context window and adjustable thinking levels. Compared with Gemini 3.6 Flash, it improves coding, tool use, multi-step planning, and instruction following.","pricing":{"cache_read":0.075,"input":0.75,"output":3.75},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"ox-alpha","model_name":"Ox Alpha","developer_id":41,"desc":"Developed by Stealth, Ox Alpha is a reasoning model designed for coding, sustained agentic work, and demanding production workloads. With a massive context length of 1,048,576 tokens, it is ideally suited for long-horizon software engineering and complex reasoning. The model excels at supporting advanced workflows that combine text with other data.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"dots-3-note-preview-free","model_name":"Dots 3 Note Preview (free)","developer_id":39,"desc":"Dots3-Note Preview is an open-weight mixture-of-experts model developed by Dots Studio, featuring 16B active parameters out of 280B total. As the lightest model in the Dots 3 family, it is designed for efficient performance while supporting an expansive context length of 512,000 tokens. This preview version provides an accessible way to experience the capabilities of the Dots 3 architecture.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":512000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.7-flash-free","model_name":"Gemini 3.7 Flash (free)","developer_id":8,"desc":"Gemini 3.7 Flash free version: fFree model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version: gemini-3.7-flash","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5.2","model_name":"GLM 5.2","developer_id":5,"desc":"GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context, execute long-running tasks more reliably, follow engineering standards more consistently, and complete the full development workflow from requirements to multi-platform deployment in a single task.","pricing":{"cache_read":0.2817,"input":1.1268,"output":3.9438},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-flash-0731","model_name":"DeepSeek V4 Flash 0731","developer_id":7,"desc":"DeepSeek-V4-Flash-0731(deepseek-v4-flash-0731) is an open-source MoE large language model developed by the Chinese AI company DeepSeek, with support for a million-token context window. It is designed for coding, complex reasoning, tool use, agentic workflows, and long-document processing. Its advantages include strong performance with fewer active parameters and improved efficiency through DSpark speculative decoding. Compared with DeepSeek V4-Flash Preview, it offers significantly stronger coding and agent capabilities, while outperforming DeepSeek V4-Pro Preview on several benchmarks with fewer active parameters.","pricing":{"cache_read":0.0284,"input":0.142,"output":0.284},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-flash-vision-exp","model_name":"DeepSeek V4 Flash Vision Exp","developer_id":7,"desc":"DeepSeek’s officially released new multimodal visual-understanding model, DeepSeek‑V4‑Flash‑Vision‑Exp, is experimental in nature and supports multimodal inputs. In pure-text capabilities (agents, reasoning, world knowledge, etc.), DeepSeek‑V4‑Flash‑Vision‑Exp is on par with the official DeepSeek‑V4‑Flash release.","pricing":{"cache_read":0.0284,"input":0.142,"output":0.284},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text,image","endpoints":"chat_completions,claude_api","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-pro-0813","model_name":"DeepSeek V4 Pro 0813","developer_id":7,"desc":"DeepSeek V4 Pro 0813 is DeepSeek’s high-performance general-purpose reasoning and agent model, designed for complex reasoning, coding, long-document analysis, and agentic workflows. It supports thinking and non-thinking modes, a 1M-token context window, up to 384K output, tool calling, and the Responses API. Compared with V4 Flash 0731, Pro prioritizes capability on complex tasks, while Flash focuses on speed, cost efficiency, and high concurrency.","pricing":{"cache_read":0.023058,"input":0.6918,"output":2.0754},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.6","model_name":"Grok 4.6","developer_id":9,"desc":"Grok 4.6 is xAI’s (SpaceXAI) flagship multimodal reasoning model for coding, long-running agents, knowledge work, and interactive application development. It supports image understanding, a 500K context window, tool calling, and structured outputs. Compared with Grok 4.5, it offers stronger multi-step execution, self-verification, coding, and visual project generation. ","pricing":{"cache_read":0.5,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":500000,"context_length":500000,"schema_checked":false,"playground_checked":false},{"model_id":"mai-thinking-1","model_name":"Mai Thinking 1","developer_id":3,"desc":"MAI-Thinking-1 is Microsoft’s first inference model in the MAI series, built for enterprise-scale workloads. With excellent reasoning, mathematical, and general intelligence capabilities, combined with superior cost-effectiveness, it makes high-throughput, 24/7 AI workloads economically viable.","pricing":{"input":2,"output":8},"types":"llm","features":"thinking,structured_outputs","input_modalities":"text","endpoints":"chat_completions","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-flash-0731-fast","model_name":"DeepSeek V4 Flash 0731 Fast","developer_id":7,"desc":"DeepSeek V4 Flash 0731 Fast is a high-speed deployment of DeepSeek’s agentic model provided by Wafer, designed for coding, tool use, and high-volume agent workloads. It preserves the capabilities of V4 Flash 0731 while delivering much faster inference. Compared with V4 Pro, it prioritizes latency and execution efficiency.","pricing":{"cache_read":0.07,"input":0.28,"output":0.56},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.6-sol-disc","model_name":"GPT 5.6 Sol Disc","developer_id":12,"desc":"GPT-5.6 Sol (limited-time 50% off) is OpenAI’s frontier reasoning model for complex coding, professional knowledge work, deep research, and long-running agents. It supports a roughly 1.05M-token context window, image understanding, and extensive tool use. Compared with Terra and Luna, Sol prioritizes capability and reliability on demanding tasks.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":30},"types":"llm","features":"tools,thinking,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false,"promotion":{"name":"gpt-5.6-sol-discount limited-time special offer","off_percent":50,"time_type":"daily","daily":{"ranges":["00:00-23:59"]}}},{"model_id":"gpt-5.6-luna","model_name":"GPT 5.6 Luna","developer_id":12,"desc":"GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads. It roughly corresponds to the nano model tier used in earlier GPT-5 families.","pricing":{"cache_read":0.02,"cache_write":0.25,"input":0.2,"output":1.2},"types":"llm","features":"tools,thinking,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.6-sol","model_name":"GPT 5.6 Sol","developer_id":12,"desc":"GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":30},"types":"llm","features":"tools,thinking,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":true},{"model_id":"gpt-5.6-terra","model_name":"GPT 5.6 Terra","developer_id":12,"desc":"GPT-5.6 Terra is designed for workloads that balance intelligence and cost. It roughly corresponds to the mini model tier used in earlier GPT-5 families.","pricing":{"cache_read":0.2,"cache_write":2.5,"input":2,"output":12},"types":"llm","features":"tools,thinking,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"agnes-2.5-flash","model_name":"Agnes 2.5 Flash","developer_id":36,"desc":"Agnes 2.5 Flash is Agnes AI’s fast and efficient language model, an upgraded, fully available model based on Agnes 2.0 Flash. It continues to use an OpenAI-compatible Chat Completions interface and has been optimized for coding tasks, agent workflows, tool calls, multi-turn dialogue, reasoning, and image understanding.","pricing":{"input":0.03,"output":0.15},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":65500,"context_length":512000,"schema_checked":false,"playground_checked":false},{"model_id":"agnes-2.5-pro","model_name":"Agnes 2.5 Pro","developer_id":36,"desc":"Agnes 2.5 Pro is Agnes AI’s paid inference model and the commercially stable version of the Agnes 2.5 Pro Alpha ranking model, suitable for advanced coding, scientific reasoning, long-context analysis, agent workflows, and multimodal understanding. The model is accessed via an OpenAI-compatible Chat Completions API.","pricing":{"cache_read":0.00378,"input":0.45,"output":0.9},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":65536,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"agnes-2.5-pro-alpha","model_name":"Agnes 2.5 Pro Alpha","developer_id":36,"desc":"Agnes 2.5 Pro Alpha is Agnes AI’s paid inference model, suitable for advanced coding, scientific reasoning, long-context analysis, agent workflows, and multimodal understanding. The model is accessed via an OpenAI-compatible Chat Completions API.","pricing":{"cache_read":0.00378,"input":0.45,"output":0.9},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":65536,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.5","model_name":"Grok 4.5","developer_id":9,"desc":"Grok 4.5 was trained on datasets spanning knowledge in coding, science, engineering, and math. With both intelligent and efficient reasoning, Grok 4.5 excels at real engineering tasks and exceeds comparable leading models at these tasks.","pricing":{"cache_read":0.5,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":500000,"context_length":500000,"schema_checked":false,"playground_checked":false},{"model_id":"lfm-2.5-2.6b-free","model_name":"Lfm 2.5 2.6b (free)","developer_id":37,"desc":"LFM-2.5-2.6B is a compact reasoning model developed by Liquid, featuring a generous 128,000 token context length. It is highly suited for agent workflows, data extraction, retrieval-augmented generation (RAG), and long-context processing. However, the developer advises against using this model for agentic coding tasks.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"muse-glimmer-30b","model_name":"Muse Glimmer 30B","developer_id":34,"desc":"Muse Glimmer 30B is a dense causal language model distilled from Muse Spark, designed for autonomous agent work. It combines multi-step reasoning, robust pattern-based tool invocation, and fault-recovery capabilities, and offers multimodal understanding via a ViT-G/14 perceptual encoder of roughly 1.8 billion parameters, supporting interleaved text and image inputs, a context window of over 131K tokens, and optional inference strength settings (from low to ultra-high). Muse Glimmer is trained on data in over 100 languages, performs strongly for its size on agent benchmarks including MCP Atlas, DeepSearch QA, Gaia2, and SWE-Bench Pro, and is released under the Apache 2.0 license.","pricing":{"cache_read":0.040005,"input":0.35,"output":1.5001},"types":"llm","features":"thinking,tools","input_modalities":"text,image","endpoints":"","max_output":131000,"context_length":131000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-lightning-3.5-30b-a3b","model_name":"Nemotron Lightning 3.5 30B A3B","developer_id":17,"desc":"NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks.\nCompared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.","pricing":{"cache_read":0.01,"input":0.05,"output":0.2},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.8-2.4t-a95b","model_name":"Qwen3.8 2.4t A95B","developer_id":13,"desc":"Qwen3.8-2.4T-A95B is Alibaba’s most powerful Qwen model to date. It is a 2.4‑trillion‑parameter sparse Mixture-of-Experts (MoE) model with approximately 95 billion active parameters. It is built for autonomous, long‑duration tasks: multi‑day code runs, reproducing research papers, and self‑improvement.","pricing":{"cache_read":0.5,"input":2,"output":6},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image","endpoints":"","max_output":262000,"context_length":262000,"schema_checked":false,"playground_checked":false},{"model_id":"ling-3.0-tiny-free","model_name":"Ling 3.0 Tiny (free)","developer_id":29,"desc":"Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, featuring 1.3B active parameters out of a total of 7.9B. It is designed for responsive agents, instruction following, and multi-turn conversations, supporting an extensive context length of 262,144 tokens. This makes it an efficient and powerful choice for conversational AI applications.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3.5-lightning-free","model_name":"Nemotron 3.5 Lightning (free)","developer_id":17,"desc":"NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks.\nCompared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.8-max","model_name":"Qwen3.8 Max","developer_id":13,"desc":"Qwen 3.8 Max(qwen3.8-max) is Alibaba Cloud’s flagship native vision-language model, built on a 2.4-trillion-parameter Mixture-of-Experts (MoE) architecture and supporting context windows of up to 1 million tokens. It is well suited for complex multimodal understanding, advanced reasoning, software development, agentic workflows, and long-context processing. At a similar price to Qwen3.7-Max, Qwen3.8-Max delivers significant improvements in reasoning, coding, and agent capabilities, with overall performance comparable to today’s leading models.","pricing":{"cache_read":0.169,"cache_write":2.1125,"input":1.69,"output":5.07},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":128000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-5","model_name":"Claude Opus 5","developer_id":2,"desc":"Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents.\n\nThe model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,structured_outputs,long_context,web","input_modalities":"text,image","endpoints":"chat_completions,claude_api,responses","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.6-flash","model_name":"Gemini 3.6 Flash","developer_id":8,"desc":"Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at code generation, agentic execution, and spatial reasoning. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.","pricing":{"cache_read":0.15,"input":1.5,"output":7.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"claude-sonnet-5","model_name":"Claude Sonnet 5","developer_id":2,"desc":"Claude Sonnet 5 is the next generation of Anthropic's Sonnet model family. It is a drop-in upgrade for Claude Sonnet 4.6 with three behavior changes: adaptive thinking is on by default, manual extended thinking now returns a 400 error (it was deprecated on Claude Sonnet 4.6), and setting sampling parameters (temperature, top_p, top_k) to non-default values returns a 400 error","pricing":{"cache_read":0.2,"cache_write":2.5,"input":2,"output":10},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,gemini_api,claude_api","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k3","model_name":"Kimi K3","developer_id":15,"desc":"Kimi K3 is Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window and industry-leading intelligence.","pricing":{"cache_read":0.3,"input":3,"output":15},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":1048576,"context_length":1048576,"schema_checked":false,"playground_checked":true},{"model_id":"ling-3.0-flash-free","model_name":"Ling 3.0 Flash (free)","developer_id":29,"desc":"Developed by Inclusionai, ling-3.0-flash-free is a 124B-parameter Mixture-of-Experts (MoE) model with approximately 5.1B parameters activated per token. Featuring an expansive context length of 262,144 tokens, this model is built to handle extensive datasets and long-form content. It is designed with token efficiency and production-scale agentic inference as key priorities to enable seamless developer deployment.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"muse-spark-1.2","model_name":"Muse Spark 1.2","developer_id":34,"desc":"Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context window.","pricing":{"input":1.375,"output":4.675},"types":"llm","features":"thinking,tools","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.8-max-preview","model_name":"Qwen3.8 Max Preview","developer_id":13,"desc":"Qwen 3.8 Max Preview(Qwen3.8-Max-Preview)  is the latest-generation foundation model in the Qwen family, packing 2.4T parameters and still evolving. Compared with the previous flagship Qwen 3.7 Max, it delivers major gains in core capabilities like Coding and Cowork (professional productivity), with world-leading performance on complex, long-horizon tasks such as full-stack development, data analysis, and Office workflows.\nLaunch offer: Credits are consumed at just 20% of the standard rate, effectively 5× your usage. Limited time only.","pricing":{"cache_read":0.0676,"cache_write":0.4225,"input":0.338,"output":1.014},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image","endpoints":"","max_output":131072,"context_length":983616,"schema_checked":false,"playground_checked":true},{"model_id":"muse-spark-1.1","model_name":"Muse Spark 1.1","developer_id":34,"desc":"Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context window.","pricing":{"input":1.375,"output":4.675},"types":"llm","features":"thinking,tools","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-flash-lite-image","model_name":"Gemini 3.1 Flash Lite Image","developer_id":8,"desc":"Google's newest, most compact, and most cost-effective image generation and editing model, designed for large-scale use.","pricing":{"cache_read":0.25,"input":0.25,"output":1.5},"types":"image_generation,llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":true,"playground_checked":false},{"model_id":"gemini-3.5-flash-lite","model_name":"Gemini 3.5 Flash Lite","developer_id":8,"desc":"Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499999},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.5-flash-lite-free","model_name":"Gemini 3.5 Flash Lite (free)","developer_id":8,"desc":"Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version:gemini-3.5-flash-lite.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.6-flash-free","model_name":"Gemini 3.6 Flash (free)","developer_id":8,"desc":"Gemini 3.6 Flash free version: fFree model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version: gemini-3.6-flash","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5.2-fast-preview","model_name":"GLM 5.2 Fast Preview","developer_id":5,"desc":"GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation.","pricing":{"cache_read":0.5635,"input":2.254,"output":7.889},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-fable-5","model_name":"Claude Fable 5","developer_id":2,"desc":"Anthropic's most capable widely released model, for the most demanding reasoning and long-horizon agentic work（This model is extremely expensive and is not recommended for casual use.）","pricing":{"cache_read":1.1,"cache_write":13.75,"input":11,"output":55},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,claude_api","max_output":128000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-8","model_name":"Claude Opus 4.8","developer_id":2,"desc":"Claude Opus 4.8 is Anthropic’s newest and most powerful publicly available model. It is suitable for the most complex tasks. It is Anthropic’s strongest model for complex reasoning, long-horizon agent programming, and highly autonomous work.claude-opus-4-8 does not display thought content by default; you need to set the additional parameter \"display\": \"summarized\" to enable it","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,claude_api","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"hy3","model_name":"Hy3","developer_id":24,"desc":"The Hy3 official version is honed for real-world business scenarios, using a Mixture-of-Experts (MoE) architecture with 295B total parameters and 21B activated parameters. It natively supports a 256K context window and offers multiple thinking modes: no_think (ultra-fast response), think_low (quick thinking), and think_high (deep reasoning), balancing ultra-fast responses, complex reasoning, and invocation cost. Compared with the Preview version, Hy3—based on real business feedback from Tencent Yuanbao, WorkBuddy, ima, Marvis, and others—focuses on improving the Coding Agent, long-form understanding, multi-turn context continuity, search QA, and complex task execution, performing more stably in reducing hallucinations, improving task completion, and engineering usability. It is better suited to practical scenarios such as frontend tasks, cross-file code development, long-document analysis, office automation, and multi-step Agent workflows.","pricing":{"cache_read":0.03905,"input":0.1562,"output":0.6248},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":128000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-1-pro","model_name":"Doubao Seed 2.1 Pro","developer_id":4,"desc":"A new generation of large models moving toward production-grade intelligence, comprehensively upgrading coding, agent, and multimodal capabilities with stronger autonomous planning, long-horizon execution, and dynamic recovery abilities to handle real, complex enterprise tasks.","pricing":{"cache_read":0.1859,"input":0.9295,"output":4.6475},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-1-turbo","model_name":"Doubao Seed 2.1 Turbo","developer_id":4,"desc":"Balancing performance and cost, comprehensively upgrading coding, agent, and multimodal capabilities, with stronger autonomous planning, long-horizon execution, and dynamic recovery abilities to handle real, complex enterprise tasks.","pricing":{"cache_read":0.09295,"input":0.46475,"output":2.32375},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"mai-image-2.5-pro","model_name":"Mai Image 2.5 Pro","developer_id":3,"desc":"MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With top-tier realism, precise text rendering, and powerful image editing capabilities, it debuted among the top three in AI image generation rankings. It is primarily targeted at commercial design, product photography, and professional creative work.","pricing":{"cache_read":5,"input":5,"output":5},"types":"image_generation,llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"grok-build-0.1","model_name":"Grok Build 0.1","developer_id":9,"desc":"Fast coding model trained specifically for agentic coding workflows.","pricing":{"cache_read":0.2,"input":1,"output":2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"mai-image-2.5","model_name":"Mai Image 2.5","developer_id":3,"desc":"MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With top-tier realism, precise text rendering, and powerful image editing capabilities, it debuted among the top three in AI image generation rankings. It is primarily targeted at commercial design, product photography, and professional creative work.","pricing":{"cache_read":5,"input":5,"output":5},"types":"image_generation,llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"mai-image-2.5-flash","model_name":"Mai Image 2.5 Flash","developer_id":3,"desc":"MAI-Image-2.5 is Microsoft's flagship AI image generation and editing model. With top-tier realism, precise text rendering, and powerful image editing capabilities, it debuted among the top three in AI image generation rankings. It is primarily targeted at commercial design, product photography, and professional creative work.","pricing":{"cache_read":1.75,"input":1.75,"output":1.75},"types":"image_generation,llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.5-flash","model_name":"Gemini 3.5 Flash","developer_id":8,"desc":"Gemini 3.5 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at sub-agent deployment, multi-step workflows, and long-horizon tasks at scale. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.","pricing":{"cache_read":0.15,"input":1.5,"output":9},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,deepsearch,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-kimi-k3","model_name":"Coding Kimi K3","developer_id":15,"desc":"","pricing":{"cache_read":0.066,"input":0.44,"output":1.61333},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":1048576,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5.2-free","model_name":"Coding GLM 5.2 (free)","developer_id":5,"desc":"coding glm 5.2 free (coding-glm-5.2-free)  is a free coding-focused API route offered by AIHubMix, powered by Z.ai’s open-source flagship GLM-5.2 model. It is suitable for code generation, debugging, project-level development, tool use, and long-running agent tasks, with support for reasoning, function calling, and structured outputs. Compared with GLM-5.1, it offers stronger coding capabilities and a 1-million-token context window. Each account is limited to 5 requests per minute, 500 requests per day, and 1 million free tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-kimi-k3-free","model_name":"Coding Kimi K3 (free)","developer_id":15,"desc":"coding-kimi-k3-free is the open and free version of coding-kimi-k3. To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":1048576,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-flash-image","model_name":"Gemini 3.1 Flash Image","developer_id":8,"desc":"gemini-3.1-flash-image (Nano Banana 2) features professional-grade visual intelligence, lightning-fast efficiency, and realistic, grounded generative capabilities. This model serves as the high-efficiency counterpart to Gemini 3 Pro Image, optimized for speed and high-volume developer use cases.","pricing":{"cache_read":0.5,"input":0.5,"output":3},"types":"image_generation,llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":true,"playground_checked":false},{"model_id":"gpt-oss-20b-free","model_name":"GPT Oss 20B (free)","developer_id":12,"desc":"Developed by OpenAI, gpt-oss-20b-free is an open-weight 21B parameter model released under the Apache 2.0 license. This model utilizes a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass. It supports an expansive context window of up to 131,072 tokens, making it well-suited for long-context tasks.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":131072,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2.7-code","model_name":"Kimi K2.7 Code","developer_id":15,"desc":"Kimi K2.7 Code is Kimi’s most intelligent Coding model, capable of completing programming tasks with higher success rates in long context. It features a native multimodal architecture that supports text, image, video input, and thinking modes, and dialogue and agent tasks.","pricing":{"cache_read":0.160835,"input":0.95,"output":3.9995},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32768,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2.7-code-highspeed","model_name":"Kimi K2.7 Code Highspeed","developer_id":15,"desc":"High-Speed version of Kimi K2.7 Code model, with output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.","pricing":{"cache_read":0.32167,"input":1.9,"output":7.999},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32768,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"longcat-2.0","model_name":"Longcat 2.0","developer_id":28,"desc":"Designed for agent development scenarios, it natively supports tool invocation, multi-step reasoning, and long-context tasks; it excels at code generation, automated workflows, and executing complex instructions; and it is deeply adapted to productivity tools such as Claude Code, OpenClaw, OpenCode, and Kilo Code.","pricing":{"cache_read":0.015492,"input":0.7746,"output":3.0984},"types":"llm","features":"thinking,tools","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-nano-9b-v2-free","model_name":"Nemotron Nano 9B V2 (free)","developer_id":17,"desc":"NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA. Designed as a unified model, it efficiently handles both reasoning and non-reasoning tasks to respond to a wide range of user queries. With a generous context length of 128,000 tokens, it is highly capable of processing long and complex documents.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3-pro-image","model_name":"Gemini 3 Pro Image","developer_id":8,"desc":"Gemini-3-Pro-Image (Nano Banana Pro) is a high-performance image generation and editing model built on Gemini 3 Pro. It delivers enhanced multimodal understanding and real-world semantic reasoning, enabling fast creation of well-structured visual content such as infographics, product sketches, and multi-subject scenes. It can also leverage real-time knowledge through Search grounding. The model excels in text rendering, consistent multi-image blending, and identity preservation, while offering fine-grained creative controls like localized edits, lighting and focus adjustments, camera transformations, and flexible aspect ratios. It’s ideal for rapid design, concept previews, product visualization, and everyday image generation workflows.","pricing":{"cache_read":2,"input":2,"output":12},"types":"image_generation,llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":true,"playground_checked":false},{"model_id":"gpt-4o-transcribe-diarize","model_name":"GPT 4o Transcribe Diarize","developer_id":12,"desc":"GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in speaker diarization, meaning it can associate audio segments in a conversation with different speakers.","pricing":{"cache_read":2.5,"input":2.5,"output":10},"types":"llm,stt","features":"thinking,function_calling,structured_outputs","input_modalities":"text,audio","endpoints":"","max_output":2000,"context_length":16000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-audio-1.5","model_name":"GPT Audio 1.5","developer_id":12,"desc":"The gpt-audio model is OpenAI's first officially released (generally available) audio model. It supports audio input and output and can be used in Chat Completions.","pricing":{"cache_read":2.5,"input":2.5,"output":10},"types":"llm,tts,stt","features":"thinking,function_calling,structured_outputs","input_modalities":"text,audio","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"hy3-preview","model_name":"Hy3 Preview","developer_id":24,"desc":"Hunyuan Hy3 preview is designed for agent workloads, adopting a MoE architecture with 295B capacity and 21B activated parameters. It provides three modes within the same model—no_think (ultra-fast response), think_low (fast thinking), and think_high (deep reasoning)—to accommodate different latency and depth requirements from high-frequency interactions to complex engineering tasks. On code benchmarks such as SWE-bench Verified it approaches the current state of the art, and its 256K context supports cross-file code refactoring and long-document analysis. It is suitable for developers who require reliable task completion while being sensitive to inference costs.","pricing":{"cache_read":0.051,"input":0.17,"output":0.566661},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":128000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m3","model_name":"MiniMax M3","developer_id":18,"desc":"The MiniMax M3 is a flagship programming model built for real-world productivity. As a production-grade model natively designed for Agent scenarios, it has achieved state-of-the-art (SOTA) performance in coding, agentic tool use, search, and office work.","pricing":{"input":0.288,"output":1.152},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":192000,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-nano-12b-v2-vl-free","model_name":"Nemotron Nano 12B V2 VL (free)","developer_id":17,"desc":"Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It features a powerful 128,000-token context length to handle extensive visual and textual inputs. The model utilizes a hybrid Transformer-Mamba architecture, successfully combining transformer-level accuracy with Mamba's structural advantages.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.7-flash","model_name":"Qwen3.7 Flash","developer_id":13,"desc":"The Qwen 3.7 series' mid-to-high cost-performance \"Plus\" model builds on strong text capabilities with a comprehensive upgrade to vision-language abilities, while retaining full agent capabilities in coding, tool use, and productivity workflows. Its core features are multimodal interactive hybrid agent capabilities, able to perceive real-world scenes, read screens and operate GUIs, generate code based on visual references, and provide end-to-end navigation of mobile applications.","pricing":{"cache_read":0.00564,"cache_write":0.03525,"input":0.0282,"output":0.1128},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.7-plus","model_name":"Qwen3.7 Plus","developer_id":13,"desc":"The Qwen 3.7 series' mid-to-high cost-performance \"Plus\" model builds on strong text capabilities with a comprehensive upgrade to vision-language abilities, while retaining full agent capabilities in coding, tool use, and productivity workflows. Its core features are multimodal interactive hybrid agent capabilities, able to perceive real-world scenes, read screens and operate GUIs, generate code based on visual references, and provide end-to-end navigation of mobile applications.","pricing":{"cache_read":0.0564,"cache_write":0.3525,"input":0.282,"output":1.128},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"step-3.7-flash","model_name":"Step 3.7 Flash","developer_id":16,"desc":"step-3.7-flash is stepfun's flagship inference model, designed for high-complexity tasks that require deep reasoning and fast execution. It excels at decomposing multi-step problems, performing tool calls, and maintaining consistency across massive datasets. It is the preferred choice for complex workloads such as long-context agents, advanced software engineering, and end-to-end research automation.","pricing":{"cache_read":0.044,"input":0.22,"output":1.32},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3-super-120b-a12b-free","model_name":"Nemotron 3 Super 120B A12B (free)","developer_id":17,"desc":"NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid Mamba-Transformer architecture. Activating just 12B parameters, it delivers maximum compute efficiency and accuracy for complex multi-agent applications. Additionally, it features an expansive context length of 262,144 tokens to handle large-scale inputs seamlessly.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-8-think","model_name":"Claude Opus 4.8 Thinking","developer_id":2,"desc":"The claude-opus-4-8-think model has adaptive thinking mode pre-enabled; the default thinking intensity is \"medium\", and it can be invoked directly via the OpenAI unified API. The claude-opus-4-8 model does not display thinking content by default; you need to set the extra parameter \"display\": \"summarized\" to enable it.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3-nano-omni-30b-a3b-reasoning-free","model_name":"Nemotron 3 Nano Omni 30B A3B (reasoning) (free)","developer_id":17,"desc":"Developed by Nvidia, NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. Supporting a massive context length of 256,000 tokens, it accepts and processes inputs across text, images, and video. This model provides powerful multimodal comprehension tailored for complex enterprise workflows.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3-ultra-550b-a55b-free","model_name":"Nemotron 3 Ultra 550B A55B (free)","developer_id":17,"desc":"NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring an extensive 1,000,000-token context length. Built on a hybrid Transformer-Mamba mixture-of-experts (MoE) architecture, it operates with 55B active parameters out of a total of 550B parameters. This advanced structure is designed to support sophisticated reasoning and complex task orchestration.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.7-max","model_name":"Qwen3.7 Max","developer_id":13,"desc":"The Max model, the largest and most capable in the Qwen3.7 series, is currently offering its pure-text model capabilities for trial. Qwen3.7 is a new-generation flagship model designed for the agent era; its core strengths lie in the breadth and depth of its agent capabilities: it performs excellently in programming, office and productivity tasks, and long-term autonomous execution.","pricing":{"cache_read":0.169,"cache_write":2.1125,"input":1.69,"output":5.07},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3.5-content-safety-free","model_name":"Nemotron 3.5 Content Safety (free)","developer_id":17,"desc":"Developed by NVIDIA, Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model fine-tuned from Google Gemma-3-4B. Supporting an expansive context length of 128,000 tokens, this model moderates both user inputs and generated responses for LLMs and VLMs. It provides robust safety filtering to ensure aligned and secure interactions across multiple modalities.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-image-2","model_name":"GPT Image 2","developer_id":12,"desc":"GPT-image-2 is OpenAI's latest cutting-edge image generation model. Key value adds include better performance, quality, editing controls, and face preservation.\n(Currently the model's image generation time may be \u003e5 minutes; it is recommended to set the client timeout to ≥10 minutes.)\nThe model supports high input_fidelity and adding/removing one aspect of the image while retaining others. This model includes improvements in aspect ratio, resolution, and editing capabilities.","pricing":{"cache_read":5,"input":5,"output":30},"types":"image_generation,llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":true,"playground_checked":false},{"model_id":"coding-glm-5.2","model_name":"Coding GLM 5.2","developer_id":5,"desc":"Currently, the special resources for this model are limited, but due to its popularity, the usage is too high, which may result in a large number of 429 errors. Resources have been coordinated, and improvements are expected in the coming weeks, so please stay tuned. In the meantime, it is recommended to use the regular-priced API model, which can ensure absolute stability. The model ID is glm-5.2.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-5.1","model_name":"ERNIE 5.1","developer_id":25,"desc":"ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its foundational capabilities and significant improvements in agents, knowledge, reasoning, and deep search. This upgrade uses a decoupled fully-asynchronous reinforcement learning technique to specifically address challenges encountered as large models evolve toward agent-based autonomous decision-making, such as training–inference numerical bias, low utilization of heterogeneous resources, and global issues caused by long-tail effects. It is paired with scaled agent post-training techniques to enhance model capabilities and generalization, enabling a three-step collaboration of environment, expert, and fusion that both ensures training efficiency and significantly improves the model’s stability and performance on complex tasks.","pricing":{"cache_read":0.5634,"input":0.5634,"output":2.5353},"types":"llm","features":"thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":119000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-flash-lite","model_name":"Gemini 3.1 Flash Lite","developer_id":8,"desc":"gemini-3.1-flash-lite is currently Google's latest and most cost-effective model, optimized for large-scale agent-based tasks, translation, and simple data processing.","pricing":{"cache_read":0.25,"input":0.25,"output":1.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,deepsearch,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-flash-lite-nothink","model_name":"Gemini 3.1 Flash Lite (no think)","developer_id":8,"desc":"gemini-3.1-flash-lite is currently Google's latest and most cost-effective model, optimized for large-scale agent-based tasks, translation, and simple data processing.","pricing":{"cache_read":0.25,"input":0.25,"output":1.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,deepsearch,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.3","model_name":"Grok 4.3","developer_id":9,"desc":"Grok 4.3 is amongst the leading models in intelligence and well priced when comparing to other models of similar price. It's also notably fast, however very verbose. The model supports text and image input, outputs text, and has a 1m tokens context window.","pricing":{"cache_read":0.2,"input":1.25,"output":2.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":1000000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"north-mini-code-free","model_name":"North Mini Code (free)","developer_id":6,"desc":"Developed by Cohere, north-mini-code-free is the debut model of the North family and Cohere's first agentic coding model. This sparse mixture-of-experts model features 30B total parameters and 3B active parameters, designed and optimized for high performance. With an expansive context length of 256,000 tokens, it is well-suited for handling complex developer workflows and large codebases.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"laguna-xs-2.1-free","model_name":"Laguna Xs 2.1 (free)","developer_id":35,"desc":"Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside, representing a step forward from the Laguna XS.2 model released in April 2026. It features an extensive context length of 262,144 tokens, allowing it to process large amounts of code and developer inputs. This model provides a robust solution for a wide range of coding agent tasks.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.5","model_name":"GPT 5.5","developer_id":12,"desc":"GPT-5.5 raises the baseline for complex production workflows. It’s a strong fit for coding use cases, tool-heavy agents, grounded assistants, long-context retrieval, product-spec-to-plan workflows, and customer-facing workflows where execution quality and response polish are critical.","pricing":{"cache_read":0.5,"input":5,"output":30},"types":"llm","features":"thinking,function_calling,structured_outputs,web,tools","input_modalities":"text,image","endpoints":"chat_completions,claude_api,responses","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.5-pro","model_name":"GPT 5.5 Pro","developer_id":12,"desc":"Please note: this model is extremely expensive and very slow. If a request fails due to network issues, you may still be charged heavily; we cannot refund charges incurred by requests to this model.\nGPT-5.5 pro is available in the Responses API only to enable support for multi-turn model interactions before responding to API requests, and other advanced API features in the future. Since GPT-5.5 pro is designed to tackle tough problems, some requests may take several minutes to finish. To avoid timeout, please set a longer timeout duration. It is recommended to use this under good network conditions.","pricing":{"cache_read":30,"input":30,"output":180},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-flash","model_name":"DeepSeek V4 Flash","developer_id":7,"desc":"(This model currently points to the older 0423 version; if you need to request the latest version, you can choose the model deepseek-v4-flash-0731)DeepSeek-V4 features an ultra-long context of one million characters and achieves leading performance domestically and in the open-source domain in agent capabilities, world knowledge, and reasoning.","pricing":{"cache_read":0.0284,"input":0.142,"output":0.284},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v4-pro","model_name":"DeepSeek V4 Pro","developer_id":7,"desc":"(This model currently points to the older 0423 version; if you need to request the latest version, you can choose the model deepseek-v4-pro-0813)DeepSeek-V4 features an ultra-long context of one million characters and achieves leading performance domestically and in the open-source domain in agent capabilities, world knowledge, and reasoning.( Directly requesting deepseek-v4-pro will route you through the official discount channel.)","pricing":{"cache_read":0.14027,"input":1.69,"output":3.38},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"","max_output":384000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemma-4-31b-it-free","model_name":"Gemma 4 31B It (free)","developer_id":8,"desc":"Gemma 4 31B Instruct is a 30.7B dense multimodal model developed by Google DeepMind that supports both text and image inputs with text outputs. It features an expansive 256K token context window alongside a configurable thinking and reasoning mode. The model also includes native function support to streamline complex tasks.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"command-a-plus-05-2026","model_name":"Command A Plus 05 2026","developer_id":6,"desc":"Cohere's stronger command model for multilingual agents and enterprise workflows","pricing":{"input":2.5,"output":10},"types":"llm","features":"thinking,tools,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,responses","max_output":64000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-5.0","model_name":"ERNIE 5.0","developer_id":25,"desc":"ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities.\n\nERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.","pricing":{"cache_read":0.82192,"input":0.82192,"output":3.28768},"types":"llm","features":"thinking","input_modalities":"text,image,audio,video","endpoints":"","max_output":64000,"context_length":119000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2.6","model_name":"Kimi K2.6","developer_id":15,"desc":"Kimi K2.6 is Kimi's latest and most intelligent model, with stronger and more stable long-range code-writing capabilities, significantly improved instruction-following and self-correction abilities, and support for text, image, and video inputs, thinking and non-thinking modes, as well as dialogue and Agent tasks.The model has a context length of 256k, supports long-form thinking, and excels at deep reasoning.","pricing":{"cache_read":0.160835,"input":0.95,"output":3.9995},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32768,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"laguna-s-2.1-free","model_name":"Laguna S 2.1 (free)","developer_id":35,"desc":"Laguna S 2.1 is the latest coding agent model from Poolside, featuring an impressive context length of 262,144 tokens. This model is built with 118B total parameters and 8B active parameters, balancing efficiency with high performance. It delivers strong capabilities for developer tasks, scoring 70.2% on the Terminal-Bench 2.1 benchmark.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-max-preview","model_name":"Qwen3.6 Max Preview","developer_id":13,"desc":"The Max model Preview version, the largest and most capable model in the Qwen3.6 series, currently offers its pure-text model capabilities for trial. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model further enhances vibe coding capabilities, executes coding agents more efficiently, significantly improves front-end programming and development capabilities, and further upgrades long-tail knowledge handling.","pricing":{"cache_read":0.1268,"cache_write":1.585,"input":1.268,"output":7.608},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":240000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2.5","model_name":"Xiaomi Mimo V2.5","developer_id":31,"desc":"MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.","pricing":{"cache_read":0.0031,"input":0.155,"output":0.31},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2.5-pro","model_name":"Xiaomi Mimo V2.5 Pro","developer_id":31,"desc":"MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent capabilities, complex software engineering, and long-horizon tasks, it can now directly compete with the world's top agent models (Claude Opus 4.6, GPT-5.4). Compared with the previous-generation MiMo-V2-Pro, it achieves an all-around leap forward.","pricing":{"cache_read":0.00384,"input":0.48,"output":0.96},"types":"llm","features":"web","input_modalities":"text","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"nemotron-3-nano-30b-a3b-free","model_name":"Nemotron 3 Nano 30B A3B (free)","developer_id":17,"desc":"NVIDIA Nemotron 3 Nano 30B A3B is a highly efficient small language Mixture of Experts (MoE) model developed by Nvidia. Designed to help developers build specialized agentic AI systems, it delivers exceptional compute efficiency and accuracy. Additionally, it features an impressive context length of 256,000 tokens to support extensive data processing.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text","endpoints":"chat_completions","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-27b","model_name":"Qwen3.6 27B","developer_id":13,"desc":"The Qwen3.6 series 27B native vision-language Dense model. Compared with the 3.5-27B, the model notably improves Agentic coding capability and further enhances STEM and reasoning abilities; on the visual modality side, spatial intelligence, object localization and detection capabilities are significantly strengthened, and video understanding, document OCR, and visual agent capabilities have steadily improved.","pricing":{"input":0.422,"output":2.532},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":254000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-35b-a3b","model_name":"Qwen3.6 35B A3B","developer_id":13,"desc":"Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding performance comparable to the current top cutting-edge models, with a significant improvement over the 3.5 series. The model's capabilities have been markedly enhanced in Agentic coding, front-end programming, Vibe coding and other coding abilities, universal multimodal recognition, OCR, object localization, and more.","pricing":{"cache_read":0.254,"input":0.254,"output":1.524},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":254000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-flash","model_name":"Qwen3.6 Flash","developer_id":13,"desc":"Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding performance comparable to the current top cutting-edge models, with a significant improvement over the 3.5 series. The model's capabilities have been markedly enhanced in Agentic coding, front-end programming, Vibe coding and other coding abilities, universal multimodal recognition, OCR, object localization, and more.","pricing":{"cache_read":0.0169,"cache_write":0.21125,"input":0.169,"output":1.014},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-7","model_name":"Claude Opus 4.7","developer_id":2,"desc":"Claude Opus 4.7 is Anthropic’s latest and most powerful publicly available model. It has high autonomy and performs exceptionally well on long-horizon agent tasks, knowledge work, vision tasks, and memory tasks.  The claude-opus-4-7 model does not display thinking content by default; you need to set the extra parameter \"display\": \"summarized\" to enable it.This page summarizes all the new features at release.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-7-think","model_name":"Claude Opus 4.7 Thinking","developer_id":2,"desc":"The claude-opus-4-7-think model has adaptive thinking mode pre-enabled; the default thinking intensity is \"medium\", and it can be invoked directly via the OpenAI unified API. The claude-opus-4-7 model does not display thinking content by default; you need to set the extra parameter \"display\": \"summarized\" to enable it.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"image,text","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-chat-latest","model_name":"GPT Chat","developer_id":12,"desc":"GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates in the future, they are routed behind this slug automatically.","pricing":{"cache_read":0.5,"input":5,"output":30},"types":"llm","features":"thinking,function_calling,structured_outputs,web,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"gemma-4-26b-a4b-it-free","model_name":"Gemma 4 26B A4B It (free)","developer_id":8,"desc":"Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model developed by Google DeepMind. Featuring an expansive context length of 262,144 tokens, it delivers near-31B quality with highly efficient inference. Despite its 25.2B total parameters, only 3.8B are activated per token, making it an incredibly fast and cost-effective solution.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"reasoning,tool_calling,long_context","input_modalities":"text,image","endpoints":"chat_completions","max_output":0,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-20-non-reasoning","model_name":"Grok 4 20","developer_id":9,"desc":"Grok 4.2 is xAI’s latest large language model, built for strong reasoning, multimodal understanding, and enterprise use. It improves instruction following, honesty, and calibration over earlier Grok versions, while supporting both single‑agent and multi‑agent workflows. Designed as a general‑purpose, truth‑seeking assistant, Grok 4.2 is well suited for research, analysis, coding, and complex professional tasks when deployed with appropriate guardrails.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-20-reasoning","model_name":"Grok 4 20 (reasoning)","developer_id":9,"desc":"Grok 4.2 is xAI’s latest large language model, built for strong reasoning, multimodal understanding, and enterprise use. It improves instruction following, honesty, and calibration over earlier Grok versions, while supporting both single‑agent and multi‑agent workflows. Designed as a general‑purpose, truth‑seeking assistant, Grok 4.2 is well suited for research, analysis, coding, and complex professional tasks when deployed with appropriate guardrails.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-plus","model_name":"Qwen3.6 Plus","developer_id":13,"desc":"Qwen 3.6, the native vision-language Plus series model, demonstrates outstanding performance comparable to the current top cutting-edge models, with a significant improvement over the 3.5 series. The model's capabilities have been markedly enhanced in Agentic coding, front-end programming, Vibe coding and other coding abilities, universal multimodal recognition, OCR, object localization, and more.","pricing":{"cache_read":0.0282,"cache_write":0.3525,"input":0.282,"output":1.692},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m3-free","model_name":"Coding MiniMax M3 (free)","developer_id":18,"desc":"coding-minimax-m3-free is a free and open version offered by AIHubMix specifically for MiniMax users.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5.1","model_name":"GLM 5.1","developer_id":5,"desc":"GLM-5.1 is Zhipu's latest flagship model, with greatly enhanced coding capabilities and significantly improved long-range task performance. It can continuously and autonomously work for up to 8 hours on a single task, completing the full closed loop from planning and execution to iterative optimization, delivering engineering-grade results.\nIn terms of general capability and coding ability, GLM-5.1's overall performance aligns with Claude Opus 4.6, and it demonstrates stronger sustained work capability in long-range autonomous execution, complex engineering optimization, and real-world development scenarios, making it an ideal foundation for building Autonomous Agents and long-horizon Coding Agents.","pricing":{"cache_read":0.183112,"input":0.845,"output":3.38},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"chat_completions,claude_api","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"gemma-4-26b-a4b-it","model_name":"Gemma 4 26B A4B It","developer_id":8,"desc":"A Mixture-of-Experts model that activates only 4B parameters per inference,delivering high-performance reasoning with a fraction of the memory cost - idealfor cost-efficient, high-throughput server deployments.","pricing":{"cache_read":0,"input":0.14,"output":0.39998},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":131100,"context_length":262100,"schema_checked":false,"playground_checked":false},{"model_id":"gemma-4-31b-it","model_name":"Gemma 4 31B It","developer_id":8,"desc":"Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license.","pricing":{"cache_read":0,"input":0.14,"output":0.39998},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":131100,"context_length":262100,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4","model_name":"GPT 5.4","developer_id":12,"desc":"GPT-5.4 is our frontier model for complex professional work.Reasoning.effort supports: none (default), low, medium, high and xhigh.","pricing":{"cache_read":0.25,"input":2.5,"output":15},"types":"llm","features":"thinking,function_calling,structured_outputs,web,tools","input_modalities":"text,image","endpoints":"chat_completions,responses,claude_api","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-sonnet-4-6","model_name":"Claude Sonnet 4.6","developer_id":2,"desc":"Claude Sonnet 4.6 delivers frontier intelligence at scale—built for coding, agents, and enterprise workflows. Use cases include: Agents: Sonnet 4.6 excels at complex, multi-step tasks requiring sustained reasoning and adaptive decision-making—ideal for workflows where reliability and autonomy matter most. Coding: Sonnet 4.6 is built for iterative development work, handling complex codebases without losing quality as you guide it through building, refactoring, and debugging. It can compress multi-day projects into hours with the technical depth to deliver production-ready solutions. Enterprise workflows: Sonnet 4.6 powers agents that manage professional projects from start to finish, leveraging memory to maintain context across files with a step-change improvement in creating spreadsheets, slides, and docs. Financial analysis: Sonnet 4.6 connects the dots across regulatory filings, market reports, and internal data—enabling sophisticated modeling and proactive compliance. Cybersecurity: Sonnet 4.6 brings professional-grade analysis to security workflows, correlating logs, vulnerability databases, and threat intelligence for proactive threat detection and automated incident response. Computer use: Sonnet 4.6 delivers confident, consistent navigation with more human-like browsing—enabling better web QA, workflow automation, and advanced user experiences.","pricing":{"cache_read":0.3,"cache_write":3.75,"input":3,"output":15},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-xiaomi-mimo-v2.5","model_name":"Coding Xiaomi Mimo V2.5","developer_id":31,"desc":"Only supports OpenAI-compatible formats.","pricing":{"cache_read":0.0016,"input":0.08,"output":0.16},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-xiaomi-mimo-v2.5-pro","model_name":"Coding Xiaomi Mimo V2.5 Pro","developer_id":31,"desc":"Only supports OpenAI-compatible formats.","pricing":{"cache_read":0.0016,"input":0.2,"output":0.4},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-lite-260428","model_name":"Doubao Seed 2.0 Lite 260428","developer_id":4,"desc":"Doubao  Coding model optimized for real-world programming environments that can reliably invoke tools in common IDEs such as Claude Code. The model is specially optimized for frontend capabilities and performs well with common frontend frameworks. The model supports Skills and can work with various custom skills.","pricing":{"cache_read":0.018082,"input":0.09041,"output":0.54246},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video,audio","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-mini-260428","model_name":"Doubao Seed 2.0 Mini 260428","developer_id":4,"desc":"Doubao  Coding model optimized for real-world programming environments that can reliably invoke tools in common IDEs such as Claude Code. The model is specially optimized for frontend capabilities and performs well with common frontend frameworks. The model supports Skills and can work with various custom skills.","pricing":{"cache_read":0.00564,"input":0.0282,"output":0.282},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-flash-image-preview","model_name":"Gemini 3.1 Flash Image Preview","developer_id":8,"desc":"gemini-3.1-flash-image-preview (Nano Banana 2) features professional-grade visual intelligence, lightning-fast efficiency, and realistic, grounded generative capabilities. This model serves as the high-efficiency counterpart to Gemini 3 Pro Image, optimized for speed and high-volume developer use cases.","pricing":{"cache_read":0.5,"input":0.5,"output":3},"types":"image_generation,llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-pro-preview","model_name":"Gemini 3.1 Pro Preview","developer_id":8,"desc":"Gemini 3.1 Pro Preview is designed to further optimize the performance and reliability of the Gemini 3 Pro series, offering improved reasoning capabilities, greater token efficiency, and a more robust, factually consistent user experience. It is optimized for software-engineering behaviors and usability, and is also suitable for agent workflows that require precise tool invocation and reliable multi-step execution, enabling stable operation across a variety of real-world scenarios.","pricing":{"cache_read":0.2,"input":2,"output":12},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,deepsearch,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-pro-preview-customtools","model_name":"Gemini 3.1 Pro Preview Customtools","developer_id":8,"desc":"gemini-3.1-pro-preview-customtools\n\nFor users who build applications mixing bash and custom tools, the Gemini 3.1 Pro preview provides a separate endpoint accessible via the API call gemini-3.1-pro-preview-customtools. This endpoint is better at prioritizing your custom tools (for example, view_file or search_code).\n\nPlease note that while gemini-3.1-pro-preview-customtools is optimized for agent workflows that use custom tools and Bash, you may experience quality fluctuations in some use cases that cannot benefit from these tools.","pricing":{"cache_read":0.2,"input":2,"output":12},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,web,deepsearch,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions,gemini_api,claude_api","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3.1-pro-preview-search","model_name":"Gemini 3.1 Pro Preview Search","developer_id":8,"desc":"Gemini-3.1-pro-preview-search integrates Google's official search functionality; the search feature incurs an additional separate fee log directly incorporated into the scoring, but the log details are not displayed; this will be fixed in the future to show the details; it only supports OpenAI-compatible format calls and does not support the Gemini SDK; for the Gemini native SDK, please directly set the official search parameters.","pricing":{"cache_read":0.2,"input":2,"output":12},"types":"llm","features":"thinking,web,deepsearch,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"chat_completions","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4-mini","model_name":"GPT 5.4 Mini","developer_id":12,"desc":"GPT-5.4 mini is a faster, more efficient model that inherits the advantages of GPT-5.4 and is specifically optimized for high-volume workloads.","pricing":{"cache_read":0.075,"input":0.75,"output":4.5},"types":"llm","features":"tools,function_calling,structured_outputs,web","input_modalities":"text,image","endpoints":"chat_completions,responses,claude_api","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4-nano","model_name":"GPT 5.4 Nano","developer_id":12,"desc":"GPT-5.4 nano is designed for tasks where speed and cost are most important, such as classification, data extraction, ranking, and sub-agent scenarios.","pricing":{"cache_read":0.02,"input":0.2,"output":1.25},"types":"llm","features":"tools,function_calling,structured_outputs,thinking,web","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.5-free","model_name":"GPT 5.5 (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\". To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,function_calling,structured_outputs,web,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-plus","model_name":"Qwen3.5 Plus","developer_id":13,"desc":"The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current leading frontier models, with leapfrog improvements over the 3 series in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15.","pricing":{"cache_read":0.01096,"cache_write":0.137,"input":0.1096,"output":0.6576},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-122b-a10b","model_name":"Qwen3.5 122B A10B","developer_id":13,"desc":"The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current leading frontier models, with leapfrog improvements over the 3 series in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15.","pricing":{"cache_read":0.1126,"input":0.1126,"output":0.9008},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-27b","model_name":"Qwen3.5 27B","developer_id":13,"desc":"The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current leading frontier models, with leapfrog improvements over the 3 series in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15.","pricing":{"cache_read":0.0846,"input":0.0846,"output":0.6768},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-35b-a3b","model_name":"Qwen3.5 35B A3B","developer_id":13,"desc":"The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current leading frontier models, with leapfrog improvements over the 3 series in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15.","pricing":{"cache_read":0.0564,"input":0.0564,"output":0.4512},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-397b-a17b","model_name":"Qwen3.5 397B A17B","developer_id":13,"desc":"The Qwen 3.5 native vision-language Plus model is built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current leading frontier models, with leapfrog improvements over the 3 series in both pure-text and multimodal capabilities. This model version is functionally equivalent to the snapshot model qwen3.5-plus-2026-02-15.","pricing":{"cache_read":0.1644,"input":0.1644,"output":0.9864},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.5-flash","model_name":"Qwen3.5 Flash","developer_id":13,"desc":"The Qwen3.5 native vision-language Flash series models are designed with a hybrid architecture that integrates linear attention mechanisms and sparse mixture-of-experts models, achieving higher inference efficiency. Compared with the 3 series, the models deliver leapfrog improvements in both pure-text and multimodal performance; they respond quickly and combine inference speed with high performance.","pricing":{"cache_read":0.00282,"cache_write":0.03525,"input":0.0282,"output":0.282},"types":"llm","features":"tools,function_calling,structured_outputs,web,long_context,thinking","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":991000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-sonnet-4-6-think","model_name":"Claude Sonnet 4.6 Thinking","developer_id":2,"desc":"Claude sonnet 4.6 does not enable reasoning mode by default. To access its deep reasoning capabilities, users would typically need to call the native Claude API. To make this capability available through an OpenAI-compatible interface, we provide the claude-sonnet-4-6-think model, which has reasoning mode pre-enabled and a default 32k-token context window, allowing it to be called directly via the OpenAI unified API.\nClaude sonnet 4.5 Think is a reasoning-focused variant of Claude sonnet 4.6 designed for advanced tasks that require rigorous reasoning, complex decision-making, and long-chain analysis. Aside from its enhanced reasoning mechanism, all other capabilities remain consistent with the standard Claude sonnet 4.6 model, making it well suited for complex engineering problem decomposition, multi-stage planning, and logic-intensive analysis.","pricing":{"cache_read":0.3,"cache_write":3.75,"input":3,"output":15},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"image,text","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-xiaomi-mimo-v2-omni","model_name":"Coding Xiaomi Mimo V2 Omni","developer_id":31,"desc":"Only supports OpenAI-compatible formats.","pricing":{"cache_read":0.016,"input":0.08,"output":0.4},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-xiaomi-mimo-v2-pro","model_name":"Coding Xiaomi Mimo V2 Pro","developer_id":31,"desc":"Only supports OpenAI-compatible formats.","pricing":{"cache_read":0.04,"input":0.2,"output":0.6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.3-chat-latest","model_name":"GPT 5.3 Chat","developer_id":12,"desc":"GPT-5.3Chat refers to the GPT-5.3 snapshot currently used in ChatGPT and is optimized for conversational use cases. While GPT-5.2 is recommended for most API applications, GPT-5.3Chat is ideal for testing the latest improvements in chat-based interactions.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.3-codex","model_name":"GPT-5.3-Codex","developer_id":12,"desc":"GPT-5.3-Codex is optimized for agentic coding tasks in Codex or similar environments. GPT-5.3-Codex supports low, medium, high, and xhigh reasoning effort settings. If you want to learn more about prompting GPT-5.3-Codex, refer to our dedicated guide.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-image-2-free","model_name":"GPT Image 2 (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\".","pricing":{"cache_read":0,"input":0,"output":0},"types":"image_generation,llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5.1","model_name":"Coding GLM 5.1","developer_id":5,"desc":"Only supports OpenAI-compatible formats.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-pro","model_name":"Doubao Seed 2.0 Pro","developer_id":4,"desc":"Doubao flagship all-purpose general model, targeting complex reasoning and long-chain task execution scenarios in the Agent era. It emphasizes multimodal understanding, long-context reasoning, structured generation, and tool-augmented execution. It excels at handling complex instructions and multi-constraint execution, reliably addressing multi-step complex planning, intricate image-text reasoning, video content understanding, and high-difficulty analysis scenarios.","pricing":{"cache_read":0.09644,"input":0.4822,"output":2.411},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":128000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4-high","model_name":"GPT 5.4 High","developer_id":12,"desc":"GPT-5.4 supports configurable reasoning effort only through the /responses endpoint. To make higher-intensity reasoning available directly via the /chat interface, GPT-5.4-High is provided as a reasoning-enhanced variant of GPT-5.4 with reasoning_effort preset to high. It is designed for tasks that require deeper analysis, stronger result consistency, and greater controllability. By applying more aggressive reasoning strategies and more effective use of extended context, the model delivers clearer and more reliable responses, making it well suited for complex agent workflows, long-chain decision-making, and reliability-critical advanced applications.","pricing":{"cache_read":0.25,"input":2.5,"output":15},"types":"llm","features":"thinking,function_calling,web,structured_outputs,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4-low","model_name":"GPT 5.4 Low","developer_id":12,"desc":"GPT-5.4 supports configuring reasoning strength only through the /responses endpoint. To make lower-overhead reasoning available directly in the /chat endpoint, the GPT-5.4-Low model is provided. This model is based on GPT-5.4 with reasoning_effort preset to low. This model is designed for use cases that are sensitive to response latency and cost. By adopting a lighter reasoning strategy, it delivers stable responses with lower latency and higher throughput. It is well suited for high-concurrency conversations, real-time interactions, basic Q\u0026A, and scenarios where deep reasoning is not required.","pricing":{"cache_read":0.25,"input":2.5,"output":15},"types":"llm","features":"thinking,function_calling,web,structured_outputs,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.4-pro","model_name":"GPT 5.4 Pro","developer_id":12,"desc":"Please note: this model is extremely expensive and very slow. If a request fails due to network issues, you may still be charged heavily; we cannot refund charges incurred by requests to this model.\nGPT-5.4 pro is available in the Responses API only to enable support for multi-turn model interactions before responding to API requests, and other advanced API features in the future. Since GPT-5.4 pro is designed to tackle tough problems, some requests may take several minutes to finish. To avoid timeout, please set a longer timeout duration. It is recommended to use this under good network conditions.","pricing":{"cache_read":30,"input":30,"output":180},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":1050000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-next","model_name":"Qwen3 Coder Next","developer_id":13,"desc":"The Qwen3 series is a next-generation code-generation model with results close to Qwen3-Coder-Plus while offering superior performance. The model is optimized for repository-level understanding, supports multi-turn tool interactions, and improves compatibility with agentic coding tools.","pricing":{"cache_read":0.137,"input":0.137,"output":0.548},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":64000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2-omni-free","model_name":"Xiaomi Mimo V2 Omni (free)","developer_id":31,"desc":"xiaomi-mimo-v2-omni-free is the open free version of xiaomi-mimo-v2-omni.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2-pro-free","model_name":"Xiaomi Mimo V2 Pro (free)","developer_id":31,"desc":"xiaomi-mimo-v2-pro-free is the open free version of xiaomi-mimo-v2-pro.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2.5-free","model_name":"Xiaomi Mimo V2.5 (free)","developer_id":31,"desc":"xiaomi-mimo-v2.5-free is the open free version of xiaomi-mimo-v2.5.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"xiaomi-mimo-v2.5-pro-free","model_name":"Xiaomi Mimo V2.5 Pro (free)","developer_id":31,"desc":"xiaomi-mimo-v2.5-pro-free is the open free version of xiaomi-mimo-v2.5-pro5.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m2.7","model_name":"MiniMax M2.7","developer_id":18,"desc":"MiniMax M2.7 can autonomously build complex Agent Harnesses and, leveraging capabilities such as Agent Teams, complex Skills, and the Tool Search tool, complete highly complex productivity tasks.","pricing":{"cache_read":0.05916,"input":0.2958,"output":1.1832},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-6","model_name":"Claude Opus 4.6","developer_id":2,"desc":"Claude Opus 4.6 is Anthropic’s latest state-of-the-art reasoning model. It features an adaptive “thinking” mode that dynamically decides when to think and how much to think. At the default effort level (high), Claude will almost always engage in thinking. At lower effort levels, it may skip thinking for simple problems.\n ⚠️ The minimum cache token for claude-opus-4-6 has been increased from 1,024 to 4,096 tokens.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5.1-free","model_name":"Coding GLM 5.1 (free)","developer_id":5,"desc":"coding-glm-5.1-free is the open and free version of coding-glm-5.1.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.7-free","model_name":"Coding MiniMax M2.7 (free)","developer_id":18,"desc":"coding-minimax-m2.7-free is a free and open version offered by AIHubMix specifically for MiniMax users.  To maintain reliable service, each account is limited to 5 requests per minute, 500 requests per day, and 1 million tokens per day.","pricing":{"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5","model_name":"GLM 5","developer_id":5,"desc":"GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.","pricing":{"cache_read":0.176,"input":0.88,"output":2.816},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":202752,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5v-turbo","model_name":"GLM 5 Vision Turbo","developer_id":5,"desc":"GLM-5V-Turbo is Zhipu's first multimodal coding foundation model, built for visual programming tasks. It can natively handle multimodal inputs such as images, videos, and text, and is adept at long-horizon planning, complex programming, and action execution; deeply adapted to Agent workflows, it can collaborate closely with agents like Claude Code and OpenClaw to complete the full closed loop of \"understand the environment → plan actions → execute tasks.\"","pricing":{"cache_read":0.169008,"input":0.7042,"output":3.09848},"types":"llm","features":"","input_modalities":"text,image,video","endpoints":"","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-6-think","model_name":"Claude Opus 4.6 Thinking","developer_id":2,"desc":"Claude Opus 4.6 does not enable reasoning mode by default. To access its deep reasoning capabilities, users would typically need to call the native Claude API. To make this capability available through an OpenAI-compatible interface, we provide the claude-opus-4-6-think model, which has reasoning mode pre-enabled and a default 32k-token context window, allowing it to be called directly via the OpenAI unified API.\nClaude Opus 4.5 Think is a reasoning-focused variant of Claude Opus 4.6 designed for advanced tasks that require rigorous reasoning, complex decision-making, and long-chain analysis. Aside from its enhanced reasoning mechanism, all other capabilities remain consistent with the standard Claude Opus 4.6 model, making it well suited for complex engineering problem decomposition, multi-stage planning, and logic-intensive analysis.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"image,text","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5-free","model_name":"Coding GLM 5 (free)","developer_id":5,"desc":"coding-glm-5-free is the open and free version of coding-glm-5. To ensure stable service performance, usage limits are in place: up to 5 requests per minute, 500 requests per day, and a daily token allowance of 1 million.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5-turbo-free","model_name":"Coding GLM 5 Turbo (free)","developer_id":5,"desc":"coding-glm-5-turbo-free is the open and free version of coding-glm-5-turbo. To ensure stable service performance, usage limits are in place: up to 5 requests per minute, 500 requests per day, and a daily token allowance of 1 million.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.5-free","model_name":"Coding MiniMax M2.5 (free)","developer_id":18,"desc":"coding-minimax-m2.5-free is a free and open version offered by AIHubMix specifically for MiniMax users. To maintain stable service operations, the following usage limits apply: a maximum of 5 requests per minute, 500 total requests per day, and a daily quota of 1 million tokens.","pricing":{"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-code-preview","model_name":"Doubao Seed 2.0 Code Preview","developer_id":4,"desc":"The Doubao 2.0 series is a coding model optimized for real programming environments, capable of reliably invoking tools in common IDEs such as Claude Code. The model is specially optimized for frontend capabilities and performs well with common frontend frameworks. The model supports using Skills and can work with various custom skills.","pricing":{"cache_read":0.09644,"input":0.4822,"output":2.411},"types":"llm","features":"thinking,web,tools,function_calling","input_modalities":"text,image,video","endpoints":"","max_output":128000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-lite-260215","model_name":"Doubao Seed 2.0 Lite 260215","developer_id":4,"desc":"Doubao  Coding model optimized for real-world programming environments that can reliably invoke tools in common IDEs such as Claude Code. The model is specially optimized for frontend capabilities and performs well with common frontend frameworks. The model supports Skills and can work with various custom skills.","pricing":{"cache_read":0.018082,"input":0.09041,"output":0.54246},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-2-0-mini","model_name":"Doubao Seed 2.0 Mini","developer_id":4,"desc":"Doubao 2.0 series is designed for low-latency, high-concurrency, and cost-sensitive scenarios, emphasizing fast responses and flexible inference deployment. Model performance is comparable to Doubao-Seed-1.6. It supports a 256k context window, four levels of thinking length, and multimodal understanding, making it suitable for lightweight tasks that prioritize cost and speed.","pricing":{"cache_read":0.006027,"input":0.030136,"output":0.30136},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3-flash-preview","model_name":"Gemini 3 Flash Preview","developer_id":8,"desc":"gemini-3-flash-preview is Google's latest released, most balanced model, excelling in speed, scale, and cutting-edge intelligence.","pricing":{"cache_read":0.05,"input":0.5,"output":3},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,audio","endpoints":"","max_output":0,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3-flash-preview-search","model_name":"Gemini 3 Flash Preview Search","developer_id":8,"desc":"Gemini-3-flash-preview-search integrates Google's official search functionality; the search feature incurs an additional separate fee log directly incorporated into the scoring, but the log details are not displayed; this will be fixed in the future to show the details; it only supports OpenAI-compatible format calls and does not support the Gemini SDK; for the Gemini native SDK, please directly set the official search parameters.","pricing":{"cache_read":0.05,"input":0.5,"output":3},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,audio","endpoints":"","max_output":1048576,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"glm-5-turbo","model_name":"GLM 5 Turbo","developer_id":5,"desc":"GLM-5-Turbo is a foundational model deeply optimized for the OpenClaw scenario. From the training stage it has been specifically optimized for the core requirements of OpenClaw tasks, enhancing key capabilities such as tool invocation, instruction following, scheduled and persistent tasks, and long-chain execution.","pricing":{"cache_read":0.24,"input":1.2,"output":3.9996},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":202752,"schema_checked":false,"playground_checked":false},{"model_id":"cc-glm-5.1","model_name":"CC GLM 5.1","developer_id":5,"desc":"Supports Claude native interface, can be directly requested in Claude Code.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-5","model_name":"Claude Opus 4.5","developer_id":2,"desc":"Claude Opus 4.5 is Anthropic’s latest frontier reasoning model, optimized for complex engineering, agentic workflows, and long-horizon computer use. It features strong multimodal capabilities, improved resistance to prompt injection, and a new Verbosity parameter to control token efficiency. With advanced tool use, extended context, and multi-agent support, Opus 4.5 excels in autonomous research, debugging, planning, and spreadsheet/browser operations.\n⚠️ The minimum cache token for claude-opus-4-5 has been increased from 1,024 to 4,096 tokens.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-5-think","model_name":"Claude Opus 4.5 Thinking","developer_id":2,"desc":"Claude Opus 4.5 does not enable reasoning mode by default. To access its deep reasoning capabilities, users would typically need to call the native Claude API. To make this capability available through an OpenAI-compatible interface, we provide the claude-opus-4-5-think model, which has reasoning mode pre-enabled and a default 32k-token context window, allowing it to be called directly via the OpenAI unified API.\nClaude Opus 4.5 Think is a reasoning-focused variant of Claude Opus 4.5 designed for advanced tasks that require rigorous reasoning, complex decision-making, and long-chain analysis. Aside from its enhanced reasoning mechanism, all other capabilities remain consistent with the standard Claude Opus 4.5 model, making it well suited for complex engineering problem decomposition, multi-stage planning, and logic-intensive analysis.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"image,text","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"mimo-v2-omni","model_name":"MiMo V2 Omni","developer_id":31,"desc":"MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind \"perception\" and \"action.\" This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.","pricing":{"cache_read":0.088,"input":0.44,"output":2.2},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"mimo-v2-pro","model_name":"MiMo V2 Pro","developer_id":31,"desc":"Xiaomi MiMo-V2-Pro is built for high-intensity agent work scenarios in the real world. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M-token context length. On top of a powerful model foundation, we continuously scale compute across broader agent scenarios, further expanding the intelligent action space and achieving significant generalization from Coding to Claw.","pricing":{"cache_read":0.22,"input":1.1,"output":3.3},"types":"llm","features":"web","input_modalities":"text,image,video,audio","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"cohere-command-a","model_name":"Cohere Command A","developer_id":6,"desc":"Command A is Cohere most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.","pricing":{"cache_read":0,"input":2.5,"output":10},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3-flash-preview-free","model_name":"Gemini 3 Flash Preview (free)","developer_id":8,"desc":"gemini-3-flash-preview-free is the free, publicly available version of gemini-3-flash-preview, offering the same model capabilities with usage limits in place to ensure service stability. Limits include up to 5 requests per minute, a maximum of 250 requests per day, and a daily quota of 500,000 tokens. Free usage is based on shared capacity and is limited in availability. This version is intended for testing and light usage; for consistent and reliable access, please switch to the paid model.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m3","model_name":"CC MiniMax M3","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m3","model_name":"Coding MiniMax M3","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1-free","model_name":"GPT 4.1 (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\".","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1-mini-free","model_name":"GPT 4.1 Mini (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\".","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1-nano-free","model_name":"GPT 4.1 Nano (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\".","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-free","model_name":"GPT 4o (free)","developer_id":12,"desc":"This free model API comes from the OpenAI model deployed on Azure. To prevent abuse, the external content filter provided by Azure has been enforced, which will result in additional delays. If you want to experience the full version of the model API without filters, please use the paid version and request the model name ID without \"-free\".","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5","model_name":"Coding GLM 5","developer_id":5,"desc":"Only supports OpenAI-compatible formats.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-5-turbo","model_name":"Coding GLM 5 Turbo","developer_id":5,"desc":"Only supports OpenAI-compatible formats.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"glm-4.7","model_name":"GLM 4.7","developer_id":5,"desc":"GLM-4.7 is Zhiyuan's latest flagship model. GLM-4.7 enhances coding capabilities, long-range task planning, and tool collaboration for Agentic Coding scenarios, achieving leading performance among open-source models on several current public benchmarks. It features improved general capabilities, with responses that are more concise and natural, and writing that is more immersive. When executing complex agent tasks and tool usage, it follows instructions more strictly, with further improvements in the frontend aesthetics of Artifacts and Agentic Coding as well as the efficiency of completing long-range tasks.","pricing":{"cache_read":0.054795,"input":0.273974,"output":1.095896},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"glm-4.7-flash-free","model_name":"GLM 4.7 Flash (free)","developer_id":5,"desc":"The glm-4.7-flash free model has usage restrictions to ensure stable service operation: a maximum of 5 requests per minute, no more than 500 requests per day, and a daily usage quota of 1 million tokens.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,structured_outputs,function_calling","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-4.7-free","model_name":"Coding GLM 4.7 (free)","developer_id":5,"desc":"coding-glm-4.7-free is the open and free version of coding-glm-4.7. To ensure stable service performance, usage limits are in place: up to 5 requests per minute, 500 requests per day, and a daily token allowance of 1 million.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-3-pro-image-preview","model_name":"Gemini 3 Pro Image Preview","developer_id":8,"desc":"Gemini-3-Pro-Image-Preview (Nano Banana Pro) is a high-performance image generation and editing model built on Gemini 3 Pro. It delivers enhanced multimodal understanding and real-world semantic reasoning, enabling fast creation of well-structured visual content such as infographics, product sketches, and multi-subject scenes. It can also leverage real-time knowledge through Search grounding. The model excels in text rendering, consistent multi-image blending, and identity preservation, while offering fine-grained creative controls like localized edits, lighting and focus adjustments, camera transformations, and flexible aspect ratios. It’s ideal for rapid design, concept previews, product visualization, and everyday image generation workflows.","pricing":{"cache_read":2,"input":2,"output":12},"types":"image_generation,llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"deepinfra-gemma-4-26b-a4b-it","model_name":"Deepinfra Gemma 4 26B A4B It","developer_id":8,"desc":"A Mixture-of-Experts model that activates only 4B parameters per inference,delivering high-performance reasoning with a fraction of the memory cost - idealfor cost-efficient, high-throughput server deployments.","pricing":{"cache_read":0.011,"input":0.088,"output":0.385},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":131100,"context_length":262100,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2-codex","model_name":"GPT-5.2-Codex","developer_id":12,"desc":"GPT-5.2-Codex is an upgraded version of GPT-5.2, optimized for agentic coding tasks in Codex and similar execution environments. The model is specifically enhanced for code generation, modification, refactoring, and automated execution workflows, enabling more efficient participation in multi-step, tool-driven programming processes. GPT-5.2-Codex supports low, medium, high, and xhigh reasoning effort settings, allowing flexible trade-offs between latency, reasoning depth, and token usage depending on task complexity.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2","model_name":"GPT 5.2","developer_id":12,"desc":"GPT-5.2 is an advanced general-purpose model that improves on GPT-5.1 with more reliable, flexible, and user-friendly interactions. It delivers clearer responses, stronger instruction-following, and adaptive reasoning that scales from simple requests to complex, multi-step tasks. With enhanced control over tone and structure and support for extended context, GPT-5.2 is well suited for agent workflows, analysis, coding, and cross-domain applications.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"thinking,function_calling,structured_outputs,web,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2-chat-latest","model_name":"GPT 5.2 Chat","developer_id":12,"desc":"GPT-5.2Chat refers to the GPT-5.2 snapshot currently used in ChatGPT and is optimized for conversational use cases. While GPT-5.2 is recommended for most API applications, GPT-5.2Chat is ideal for testing the latest improvements in chat-based interactions.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2-high","model_name":"GPT 5.2 High","developer_id":12,"desc":"GPT-5.2 supports configurable reasoning effort only through the /responses endpoint. To make higher-intensity reasoning available directly via the /chat interface, GPT-5.2-High is provided as a reasoning-enhanced variant of GPT-5.2 with reasoning_effort preset to high. It is designed for tasks that require deeper analysis, stronger result consistency, and greater controllability. By applying more aggressive reasoning strategies and more effective use of extended context, the model delivers clearer and more reliable responses, making it well suited for complex agent workflows, long-chain decision-making, and reliability-critical advanced applications.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"thinking,function_calling,web,structured_outputs,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2-low","model_name":"GPT 5.2 Low","developer_id":12,"desc":"GPT-5.2 supports configuring reasoning strength only through the /responses endpoint. To make lower-overhead reasoning available directly in the /chat endpoint, the GPT-5.2-Low model is provided. This model is based on GPT-5.2 with reasoning_effort preset to low. This model is designed for use cases that are sensitive to response latency and cost. By adopting a lighter reasoning strategy, it delivers stable responses with lower latency and higher throughput. It is well suited for high-concurrency conversations, real-time interactions, basic Q\u0026A, and scenarios where deep reasoning is not required.","pricing":{"cache_read":0.175,"input":1.75,"output":14},"types":"llm","features":"thinking,function_calling,web,structured_outputs,tools","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.2-pro","model_name":"GPT 5.2 Pro","developer_id":12,"desc":"GPT-5.2 pro is available in the Responses API only to enable support for multi-turn model interactions before responding to API requests, and other advanced API features in the future. Since GPT-5.2 pro is designed to tackle tough problems, some requests may take several minutes to finish. To avoid timeout, please set a longer timeout duration. It is recommended to use this under good network conditions.","pricing":{"cache_read":2.1,"input":21,"output":168},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.1","model_name":"GPT 5.1","developer_id":12,"desc":"GPT-5 is OpenAI’s most advanced language model, designed for complex tasks that require step-by-step reasoning, precise instruction following, and high reliability. It improves reasoning, code generation, and prompt understanding—including test-time routing and intent cues like “think hard about this”—while reducing hallucination and sycophancy.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,web,tools,deepsearch,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.1-codex-max","model_name":"GPT-5.1-Codex Max","developer_id":12,"desc":"GPT-5.1-Codex-Max is a frontier programming model built for the agent-driven era. Powered by an upgraded core reasoning architecture, it is specially trained for complex agentic tasks in software engineering, mathematics, and scientific research. It delivers faster performance, greater stability, and higher token efficiency across the entire development lifecycle, including code generation, refactoring, debugging, and engineering collaboration. With native support for multiple context windows and a built-in compaction mechanism, the model can coherently process millions of tokens within a single task.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"function_calling,structured_outputs,thinking","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-1-8","model_name":"Doubao Seed 1.8","developer_id":4,"desc":"Doubao's strongest multimodal Agent model Seed1.8 has powerful multimodal capabilities, supports image and text input, and can efficiently and accurately complete tasks in scenarios such as information retrieval, code generation, GUI interaction, and complex workflows, meeting increasingly diverse technical demands.","pricing":{"cache_read":0.021918,"input":0.10959,"output":0.273975},"types":"llm","features":"thinking,web,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":64000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.1-chat-latest","model_name":"GPT 5.1 Chat","developer_id":12,"desc":"GPT-5.1 Chat refers to the GPT-5.1 snapshot currently used in ChatGPT and is optimized for conversational use cases. While GPT-5.1 is recommended for most API applications, GPT-5.1 Chat is ideal for testing the latest improvements in chat-based interactions.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.1-codex","model_name":"GPT-5.1-Codex","developer_id":12,"desc":"GPT-5.1-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments. It's available in the Responses API only and the underlying model snapshot will be regularly updated. ","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5.1-codex-mini","model_name":"GPT-5.1-Codex Mini","developer_id":12,"desc":"GPT-5.1 Codex mini is a smaller, more cost-effective, less-capable version of GPT-5.1-Codex.","pricing":{"cache_read":0.025,"input":0.25,"output":2},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-haiku-4-5","model_name":"Claude Haiku 4.5","developer_id":2,"desc":"Claude Haiku 4.5 is a fast, affordable, and highly capable AI model, excelling at coding and agentic tasks. Its combination of speed and low cost makes it ideal for powering real-time applications like chatbots, high-volume free services, and specialized \"sub-agents\" for complex tasks in coding, finance, and research. It can also handle common business tasks like creating office documents and assisting with strategy and analysis.\n⚠️ The minimum cache token for claude-haiku-4-5 has been increased from 1,024 to 4,096 tokens.","pricing":{"cache_read":0.11,"cache_write":1.375,"input":1.1,"output":5.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":131072,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"claude-sonnet-4-5","model_name":"Claude Sonnet 4.5","developer_id":2,"desc":"Sonnet 4.5 is the best model in the world for agents, coding, and computer usage. It is also our most accurate and detailed model for long-running tasks, with enhanced knowledge in coding, finance, and cybersecurity.  \nThis model supports a thinking parameter to enable thinking requests in Claude mode.","pricing":{"cache_read":0.33,"cache_write":4.125,"input":3.3,"output":16.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-sonnet-4-5-think","model_name":"Claude Sonnet 4.5 Thinking","developer_id":2,"desc":"Claude Sonnet 4.5 does not enable reasoning mode by default. To access its deep reasoning capabilities, users would typically need to call the native Claude API. To make this capability available through an OpenAI-compatible interface, the claude-sonnet-4-5-think model is provided with reasoning mode pre-enabled and a default 64k-token context window, allowing it to be called directly via the OpenAI unified API. Claude Sonnet 4.5 Think is a reasoning-focused variant of Claude Sonnet 4.5 designed for advanced tasks that require rigorous reasoning, complex decision-making, and long-chain analysis; aside from its enhanced reasoning mechanism, all other capabilities remain consistent with the standard Claude Sonnet 4.5 model, making it well suited for complex problem decomposition, multi-step planning, and logic-intensive analysis.","pricing":{"cache_read":0.33,"cache_write":4.125,"input":3.3,"output":16.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":64000,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.20-multi-agent-0309","model_name":"Grok 4.20 Multi Agent 0309","developer_id":9,"desc":"Grok 4.20 is our newest flagship model with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"mistral-large-3","model_name":"Mistral Large 3","developer_id":10,"desc":"Mistral Large 3 is a MoE model with 67.5B total parameters and 41B active parameters, supporting a 256K-token context window. Trained from scratch on 3,000 NVIDIA H200 GPUs, it is one of the strongest permissively licensed open-weight models available.\n\nDesigned for advanced reasoning and long-context understanding, Mistral Large 3 delivers performance on par with the best instruction-tuned open-weight models for general-purpose tasks, while also offering image understanding capabilities. Its multilingual strengths are particularly notable for non-English/Chinese languages, making it well-suited for global applications.\n\nTypical use cases include enterprise assistants, multilingual customer support, content generation and editing, data analysis over long documents, code assistance, and research workflows that require handling large corpora or complex instructions. With its MoE architecture, Mistral Large 3 balances strong performance with efficient inference, providing a versatile backbone for building reliable, production-grade AI systems.","pricing":{"input":0.5,"output":1.5},"types":"llm","features":"function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-1-fast-non-reasoning","model_name":"Grok 4.1 Fast","developer_id":9,"desc":"Grok 4.1 is a new conversational model with significant improvements in real-world usability, delivering exceptional performance in creative, emotional, and collaborative interactions. It is more perceptive to nuanced user intent, more engaging to converse with, and more coherent in personality, while fully preserving its core intelligence and reliability. Built on large-scale reinforcement learning infrastructure, the model is optimized for style, personality, helpfulness, and alignment, and leverages frontier agentic reasoning models as reward evaluators to autonomously assess and iterate on responses at scale, significantly enhancing overall interaction quality.","pricing":{"cache_read":0.05,"input":0.2,"output":0.5},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-1-fast-reasoning","model_name":"Grok 4.1 Fast (reasoning)","developer_id":9,"desc":"Grok 4.1 is a new conversational model with significant improvements in real-world usability, delivering exceptional performance in creative, emotional, and collaborative interactions. It is more perceptive to nuanced user intent, more engaging to converse with, and more coherent in personality, while fully preserving its core intelligence and reliability. Built on large-scale reinforcement learning infrastructure, the model is optimized for style, personality, helpfulness, and alignment, and leverages frontier agentic reasoning models as reward evaluators to autonomously assess and iterate on responses at scale, significantly enhancing overall interaction quality.","pricing":{"cache_read":0.05,"input":0.2,"output":0.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-code-fast-1","model_name":"Grok Code Fast 1","developer_id":9,"desc":"Grok 4.1 is a new conversational model with significant improvements in real-world usability, delivering exceptional performance in creative, emotional, and collaborative interactions. It is more perceptive to nuanced user intent, more engaging to converse with, and more coherent in personality, while fully preserving its core intelligence and reliability. Built on large-scale reinforcement learning infrastructure, the model is optimized for style, personality, helpfulness, and alignment, and leverages frontier agentic reasoning models as reward evaluators to autonomously assess and iterate on responses at scale, significantly enhancing overall interaction quality.","pricing":{"cache_read":0.05,"input":0.2,"output":0.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":10000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"k2.6-code-preview-free","model_name":"K2.6 Code Preview (free)","developer_id":15,"desc":"kimi-for-coding-free is a free and open version offered by AIHubMix specifically for Kimi users. To maintain stable service operations, the following usage limits apply: a maximum of 5 requests per minute 500 total requests per day, and a daily quota of 1 million tokens.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"mimo-v2-flash","model_name":"MiMo V2 Flash","developer_id":31,"desc":"MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.","pricing":{"cache_read":0.03836,"input":0.1918,"output":0.5754},"types":"llm","features":"web","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3.6-plus-preview-free","model_name":"Qwen3.6 Plus Preview (free)","developer_id":13,"desc":"This model has been removed from the platform.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,long_context","input_modalities":"text","endpoints":"","max_output":65535,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"zai-glm-5-turbo","model_name":"Zai Glm 5 Turbo","developer_id":5,"desc":"","pricing":{"cache_read":0.24,"input":1.2,"output":3.9996},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-glm-5","model_name":"CC GLM 5","developer_id":5,"desc":"Supports Claude native interface, can be directly requested in Claude Code.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-glm-5-turbo","model_name":"CC GLM 5 Turbo","developer_id":5,"desc":"Supports Claude native interface, can be directly requested in Claude Code.","pricing":{"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cloudflare-glm-5.2","model_name":"Cloudflare Glm 5.2","developer_id":5,"desc":"","pricing":{"cache_read":0.2604,"input":1.4,"output":4.4002},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-image","model_name":"Gemini 2.5 Flash Image","developer_id":8,"desc":"Gemini 2.5 Flash Image (Nano-Banana) is a state-of-the-art image generation and editing model that enables seamless blending of multiple images into a single composition while maintaining character consistency for rich visual storytelling. It supports precise, targeted image transformations through natural language instructions and leverages built-in world knowledge for both image generation and editing, making it well suited for creative design, content production, advertising, and visual expression workflows.","pricing":{"cache_read":0.3,"input":0.3,"output":2.499},"types":"image_generation,llm","features":"","input_modalities":"image,text","endpoints":"","max_output":8000,"context_length":32800,"schema_checked":true,"playground_checked":false},{"model_id":"gpt-5","model_name":"GPT 5","developer_id":12,"desc":"GPT-5 is OpenAI’s most advanced general-purpose model, delivering major improvements in reasoning, code quality, and overall user experience. It is optimized for complex tasks that require step-by-step reasoning, precise instruction following, and high accuracy in high-stakes scenarios. The model supports test-time routing and advanced prompt understanding, including user-specified intent such as “think hard about this,” while significantly reducing hallucination and sycophancy and improving performance in coding, writing, and health-related tasks.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v3.2","model_name":"DeepSeek V3.2","developer_id":7,"desc":"DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse Attention and reinforced reasoning performance, but its core strength lies in powerful agentic capabilities—enabled by large-scale task-synthesis that tightly integrates reasoning with real-world tool use, delivering robust, compliant, and generalizable agent behaviour. Users can toggle deeper reasoning through the reasoning_enabled switch.","pricing":{"cache_read":0.0302,"input":0.302,"output":0.453},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-v3.2-think","model_name":"DeepSeek V3.2 Thinking","developer_id":7,"desc":"DeepSeek-V3.2 is an efficient large language model equipped with DeepSeek Sparse Attention and reinforced reasoning performance, but its core strength lies in powerful agentic capabilities—enabled by large-scale task-synthesis that tightly integrates reasoning with real-world tool use, delivering robust, compliant, and generalizable agent behaviour. Users can toggle deeper reasoning through the reasoning_enabled switch.","pricing":{"cache_read":0.0302,"input":0.302,"output":0.453},"types":"llm","features":"tools,function_calling,structured_outputs,thinking","input_modalities":"text","endpoints":"","max_output":64000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5-codex","model_name":"GPT-5-Codex","developer_id":12,"desc":"GPT-5-Codex is a version of GPT-5 optimized for autonomous coding tasks in Codex or similar environments. It is only available in the Responses API, and the underlying model snapshots will be updated regularly. https://docs.aihubmix.com/en/api/Responses-API You can also use it in codex-cll; see https://docs.aihubmix.com/en/api/Codex-CLI for using codex-cll through Aihubmix.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-V3.1-Terminus","model_name":"DeepSeek V3.1 Terminus","developer_id":7,"desc":"DeepSeek-V3.1 non-thinking mode has now been updated to the DeepSeek-V3.1-Terminus version.","pricing":{"input":0.56,"output":1.68},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":32000,"context_length":160000,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-V3.1-Think","model_name":"DeepSeek V3.1 Thinking","developer_id":7,"desc":"Thinking mode of DeepSeek-V3.1;  \nDeepSeek V3.1 is a text generation model provided by DeepSeek, featuring a hybrid reasoning architecture that achieves an effective integration of thinking and non-thinking modes.","pricing":{"input":0.56,"output":1.68},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":32000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5-pro","model_name":"GPT 5 Pro","developer_id":12,"desc":"GPT-5 pro uses more compute to think harder and provide consistently better answers.\n\nGPT-5 pro is available in the Responses API only to enable support for multi-turn model interactions before responding to API requests, and other advanced API features in the future. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish. To avoid timeouts, try using background mode. As our most advanced reasoning model, GPT-5 pro defaults to (and only supports) reasoning.effort: high. GPT-5 pro does not support code interpreter.","pricing":{"input":15,"output":120},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5-mini","model_name":"GPT 5 Mini","developer_id":12,"desc":"GPT-5 mini is a faster, more cost-efficient version of GPT-5. It's great for well-defined tasks and precise prompts.","pricing":{"cache_read":0.025,"input":0.25,"output":2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5-nano","model_name":"GPT 5 Nano","developer_id":12,"desc":"GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, designed specifically for developer tools and environments that demand rapid interactions and ultra-low latency. While it offers a more lightweight solution with limited reasoning depth compared to its larger counterparts, GPT-5-Nano excels in core capabilities such as instruction-following and maintaining critical safety features. As the successor to GPT-4.1-nano, it provides an optimal choice for cost-sensitive or real-time applications, where efficiency and speed are paramount. Particularly well-suited for summarization and classification tasks, GPT-5-Nano is a powerful tool for developers needing a swift, reliable AI model for streamlined processes.","pricing":{"cache_read":0.005,"input":0.05,"output":0.4},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-5-chat-latest","model_name":"GPT 5 Chat","developer_id":12,"desc":"GPT-5 Chat points to the GPT-5 snapshot currently used in ChatGPT. GPT-5 is our next-generation, high-intelligence flagship model. It accepts both text and image inputs, and produces text outputs.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":400000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-opus-4-1","model_name":"Claude Opus 4.1","developer_id":2,"desc":"Opus 4.1 is an upgraded version of Claude Opus 4, with improvements mainly in agent tasks, practical coding, and reasoning. Compared to Opus 4, there is a slight improvement in software engineering accuracy; Opus 4.1 has higher accuracy at 74.5%.","pricing":{"input":16.5,"output":82.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2.5","model_name":"Kimi K2.5","developer_id":15,"desc":"Kimi K2.5 is the smartest model of Kimi to date, achieving open-source state-of-the-art (SoTA) performance in Agent, coding, visual understanding, and a series of general intelligent tasks. At the same time, Kimi K2.5 is also the most versatile model of Kimi so far, with a native multimodal architecture design that supports both visual and text input, thinking and non-thinking modes, as well as dialogue and Agent tasks.","pricing":{"cache_read":0.105,"input":0.6,"output":3},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-max-2026-01-23","model_name":"Qwen3 Max 2026 01-23","developer_id":13,"desc":"The snapshot version of the Tongyi Qianwen 3 series Max model is from January 23, 2026. By default, it does not require thinking, but thinking mode can be enabled through the enable_thinking parameter, as detailed in the code example. (After enabling thinking by passing parameters, it becomes: Qwen3-Max-Thinking). This model has a total parameter count exceeding one trillion (1T) and a pre-training data volume of up to 36T Tokens, making it the largest and most powerful reasoning model from Alibaba to date.","pricing":{"cache_read":0.09016,"cache_write":0.5635,"input":0.4508,"output":1.8032},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":32000,"context_length":252000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-flash","model_name":"Qwen3 VL Flash","developer_id":13,"desc":"The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.","pricing":{"cache_read":0.00412,"input":0.0206,"output":0.206},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":254000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-flash-2026-01-22","model_name":"Qwen3 VL Flash 2026 01-22","developer_id":13,"desc":"The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.","pricing":{"cache_read":0.0206,"input":0.0206,"output":0.206},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":254000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-plus","model_name":"Qwen3 VL Plus","developer_id":13,"desc":"The Qwen3 series visual understanding model achieves an effective fusion of thinking and non-thinking modes. Its visual agent capabilities reach world-class levels on public test sets such as OS World. This version features comprehensive upgrades in visual coding, spatial perception, and multimodal reasoning; visual perception and recognition abilities are greatly enhanced, supporting ultra-long video understanding.","pricing":{"cache_read":0.0274,"input":0.137,"output":1.37},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m2.5","model_name":"MiniMax M2.5","developer_id":18,"desc":"The MiniMax M2.5 is a flagship programming model built for real-world productivity. As a production-grade model natively designed for Agent scenarios, it has achieved state-of-the-art (SOTA) performance in coding, agentic tool use, search, and office work.","pricing":{"input":0.288,"output":1.152},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":192000,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m2.5-highspeed","model_name":"MiniMax M2.5 Highspeed","developer_id":18,"desc":"• Same performance as minimax-m2.5\n• Significantly faster inference","pricing":{"input":0.288,"output":1.152},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":192000,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"mm-minimax-m2.7-highspeed","model_name":"Mm Minimax M2.7 Highspeed","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2.7","model_name":"CC MiniMax M2.7","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2.7-highspeed","model_name":"CC MiniMax M2.7 Highspeed","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.7","model_name":"Coding MiniMax M2.7","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.7-highspeed","model_name":"Coding MiniMax M2.7 Highspeed","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2.5","model_name":"CC MiniMax M2.5","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2.5-highspeed","model_name":"CC MiniMax M2.5 Highspeed","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.5","model_name":"Coding MiniMax M2.5","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.5-highspeed","model_name":"Coding MiniMax M2.5 Highspeed","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2.1","model_name":"CC MiniMax M2.1","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-4.7","model_name":"Coding GLM 4.7","developer_id":5,"desc":"Only supports OpenAI-compatible formats.","pricing":{"cache_read":0.010998,"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.1","model_name":"Coding MiniMax M2.1","developer_id":18,"desc":"","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2.1-free","model_name":"Coding MiniMax M2.1 (free)","developer_id":18,"desc":"coding-minimax-m2.1-free is a free and open version offered by AIHubMix specifically for MiniMax users. To maintain stable service operations, the following usage limits apply: a maximum of 5 requests per minute, 500 total requests per day, and a daily quota of 1 million tokens.","pricing":{"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-audio-preview","model_name":"GPT 4o Audio Preview","developer_id":12,"desc":"OpenAI voice input and output model, with prices consistent with the official ones. For now, only the text portion prices are displayed; voice prices can be found on the official OpenAI website. Backend billing is the same as the official.","pricing":{"input":2.5,"output":10},"types":"llm","features":"","input_modalities":"text,audio","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-mini-audio-preview","model_name":"GPT 4o Mini Audio Preview","developer_id":12,"desc":"","pricing":{"input":0.15,"output":0.6},"types":"llm","features":"","input_modalities":"text,audio","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m2.1","model_name":"MiniMax M2.1","developer_id":18,"desc":"MiniMax-M2.1 redefines efficiency for intelligent agents. It is a compact, fast, and cost-effective MoE model with a total of 230 billion parameters and 10 billion active parameters, designed for top performance in coding and intelligent agent tasks while maintaining strong general intelligence. With only 10 billion active parameters, MiniMax-M2 delivers the complex end-to-end tool usage performance expected from today's leading models, but in a more streamlined form factor, making deployment and scaling easier than ever before.","pricing":{"input":0.288,"output":1.152},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":192000,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"o3","model_name":"O3","developer_id":12,"desc":"OpenAI o3 is a powerful model across multiple domains, setting a new standard for coding, math, science, and visual reasoning tasks.","pricing":{"cache_read":0.5,"input":2,"output":8},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":100000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-for-coding-free","model_name":"Kimi For Coding (free)","developer_id":15,"desc":"kimi-for-coding-free is a free and open version offered by AIHubMix specifically for Kimi users. To maintain stable service operations, the following usage limits apply: a maximum of 5 requests per minute 500 total requests per day, and a daily quota of 1 million tokens.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"o3-pro","model_name":"O3 Pro","developer_id":12,"desc":"o3-pro\nThis model only supports Requests API interface requests.The model's thinking time is relatively long, so the response will be slow.","pricing":{"cache_read":20,"input":20,"output":80},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":100000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"step-3.5-flash","model_name":"Step 3.5 Flash","developer_id":16,"desc":"step-3.5-flash is stepfun's flagship inference model, designed for high-complexity tasks that require deep reasoning and fast execution. It excels at decomposing multi-step problems, performing tool calls, and maintaining consistency across massive datasets. It is the preferred choice for complex workloads such as long-context agents, advanced software engineering, and end-to-end research automation.","pricing":{"input":0.11,"output":0.33},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-4.6","model_name":"Coding GLM 4.6","developer_id":5,"desc":"","pricing":{"cache_read":0.010998,"input":0.06,"output":0.22},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"coding-glm-4.6-free","model_name":"Coding GLM 4.6 (free)","developer_id":5,"desc":"coding-glm-4.6-free is the open and free version of coding-glm-4.6. To ensure stable service performance, usage limits are in place: up to 5 requests per minute, 500 requests per day, and a daily token allowance of 1 million.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2","model_name":"Coding MiniMax M2","developer_id":18,"desc":"coding-minimax-m2 is a free and open version offered by AIHubMix specifically for MiniMax users. To maintain stable service operations, the following usage limits apply: a maximum of 10 requests per minute, 1,000 total requests per day, and a daily quota of 5 million tokens.204800","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"coding-minimax-m2-free","model_name":"Coding MiniMax M2 (free)","developer_id":18,"desc":"coding-minimax-m2-free is a free and open version offered by AIHubMix specifically for MiniMax users. To maintain stable service operations, the following usage limits apply: a maximum of 5 requests per minute, 500 total requests per day, and a daily quota of 1 million tokens.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":13100,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro","model_name":"Gemini 2.5 Pro","developer_id":8,"desc":"Gemini 2.5 Pro is an advanced reasoning model developed by Google, optimized for solving highly complex problems across multiple domains. It can deeply understand large-scale information from diverse sources, including text, audio, images, video, and even entire codebases. The model demonstrates strong reasoning capabilities in coding, mathematics, and STEM-related tasks, and supports long-context analysis for large datasets, codebases, and technical documentation.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"tools,function_calling,structured_outputs,long_context,web,thinking,deepsearch","input_modalities":"text,image,audio,video,pdf","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"glm-4.6","model_name":"GLM 4.6","developer_id":5,"desc":"GLM-4.6 is Zhipu’s latest flagship model (total parameters 355B, activation parameters 32B), comprehensively surpassing GLM-4.5. Its coding capability is aligned with Claude Sonnet 4, making it a top domestic coding model; the context window has been expanded from 128K to 200K, better suited for long code and agent tasks; inference capabilities have been significantly enhanced and support tool invocation during processing; improvements have been made in tool calling, search agents, writing style, role play, and multilingual translation. The model is named glm-4.6 and is provided by three vendors, with calls prioritized to the Sophnet platform.","pricing":{"cache_read":0.054795,"input":0.273974,"output":1.095896},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":131072,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"glm-4.6v","model_name":"GLM 4.6 Vision","developer_id":5,"desc":"Zhipu's latest visual reasoning model achieves state-of-the-art visual understanding accuracy at the same scale upon release. It natively supports tool invocation, can automatically complete tasks, supports ultra-long 128K context length, and allows flexible toggling of reasoning.","pricing":{"cache_read":0.0274,"input":0.137,"output":0.411},"types":"llm","features":"","input_modalities":"text,image,video","endpoints":"","max_output":128000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-search","model_name":"Gemini 2.5 Pro Search","developer_id":8,"desc":"gemini-2.5-pro-search integrates Google's official search functionality; the search feature will have an additional separate fee log directly incorporated into the scoring, with detailed logs not displayed; this will be fixed and displayed later; only supports OpenAI-compatible formats for invocation, does not support Gemini SDK; for Gemini's native SDK, please set parameters directly using the official search parameters.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm,search","features":"thinking,web,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video,pdf","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2-thinking","model_name":"Kimi K2 Thinking","developer_id":15,"desc":"Kimi K2 Thinking is Moonshot AI's most advanced open-source inference model to date, extending the K2 series into intelligent agent and long-context inference domains. The model is built on the trillion-parameter mixture of experts (MoE) architecture introduced by Kimi K2, activating 32 billion parameters per forward pass and supporting a context window of 256,000 tokens.","pricing":{"cache_read":0.137,"input":0.548,"output":2.192},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":262144,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash","model_name":"Gemini 2.5 Flash","developer_id":8,"desc":"Gemini 2.5 Flash is Google’s best model in terms of both performance and cost efficiency, offering a comprehensive set of capabilities. It is the first Flash model to support visible reasoning, allowing insight into the thought process behind its responses. With its strong price–performance ratio, the model is well suited for large-scale processing, low-latency, high-throughput tasks that require reasoning, as well as agent-based application scenarios.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-preview-09-2025","model_name":"Gemini 2.5 Flash Preview 09 2025","developer_id":8,"desc":"This latest 2.5 Flash model comes with improvements in two key areas we heard consistent feedback on:\n\nBetter agentic tool use: We've improved how the model uses tools, leading to better performance in more complex, agentic and multi-step applications. This model shows noticeable improvements on key agentic benchmarks, including a 5% gain on SWE-Bench Verified, compared to our last release (48.9% → 54%). More efficient: With thinking on, the model is now significantly more cost-efficient—achieving higher quality outputs while using fewer tokens, reducing latency and cost (see charts above).","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"glm-4.5v","model_name":"GLM 4.5 Vision","developer_id":5,"desc":"GLM-4.5V is a vision-language foundational model designed for multimodal agent applications. Based on a mixture-of-experts (MoE) architecture, it has 106 billion parameters and 12 billion active parameters. It delivers outstanding performance in video understanding, image question answering, OCR, and document parsing, and achieves significant improvements in front-end web encoding, basic reasoning, and spatial reasoning.","pricing":{"cache_read":0.274,"input":0.274,"output":0.822},"types":"llm,ocr","features":"","input_modalities":"text,image,video","endpoints":"","max_output":16384,"context_length":64000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-lite","model_name":"Gemini 2.5 Flash Lite","developer_id":8,"desc":"Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.","pricing":{"cache_read":0.01,"input":0.1,"output":0.4},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-lite-nothink","model_name":"Gemini 2.5 Flash Lite (no think)","developer_id":8,"desc":"Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.","pricing":{"cache_read":0.01,"input":0.1,"output":0.4},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-lite-preview-09-2025","model_name":"Gemini 2.5 Flash Lite Preview 09 2025","developer_id":8,"desc":"gemini-2.5-flash-lite latest preview version","pricing":{"cache_read":0.01,"input":0.1,"output":0.4},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-lite-preview-09-2025-nothink","model_name":"Gemini 2.5 Flash Lite Preview 09 2025 (no think)","developer_id":8,"desc":"gemini-2.5-flash-lite latest preview version","pricing":{"cache_read":0.01,"input":0.1,"output":0.4},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-nothink","model_name":"Gemini 2.5 Flash (no think)","developer_id":8,"desc":"Gemini-2.5-flash defaults to thinking enabled; to disable thinking, request the name gemini-2.5-flash-nothink, which only supports OpenAI-compatible format calls and does not support Gemini SDK; for the native Gemini SDK, please set the parameter budget=0 directly.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-search","model_name":"Gemini 2.5 Flash Search","developer_id":8,"desc":"gemini-2.5-flash-search integrates Google's official search functionality; the search feature will have an additional separate fee log directly incorporated into the scoring, with detailed logs not displayed; this will be fixed and displayed later; only supports OpenAI-compatible formats for invocation, does not support Gemini SDK; for Gemini's native SDK, please set parameters directly using the official search parameters.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm,search","features":"web,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-preview-05-20-nothink","model_name":"Gemini 2.5 Flash Preview 05-20 (no think)","developer_id":8,"desc":"Gemini-2.5-flash-preview-05-20 is enabled by default for thinking; to disable it, request the name gemini-2.5-flash-preview-05-20-nothink.Only OpenAI-compatible format calls are supported; Gemini SDK is not supported. For the native Gemini SDK, please set the parameter budget=0 directly.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-flash-preview-05-20-search","model_name":"Gemini 2.5 Flash Preview 05-20 Search","developer_id":8,"desc":"Gemini-2.5 Flash Preview 05-20 Search integrates Google's official search functionality; the search feature will have an additional separate fee log directly integrated into the scoring deduction, with detailed logs not displayed. It will be fixed and displayed later. Only OpenAI-compatible formats are supported for invocation; Gemini SDK is not supported. For Gemini's native SDK, please set parameters directly using the official search parameters.","pricing":{"cache_read":0.03,"input":0.3,"output":2.499},"types":"llm,search","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-V3-Fast","model_name":"DeepSeek V3 Fast","developer_id":7,"desc":"V3 Ultra-Fast Version,The current price is a limited-time 50% discount and will return to the original price on July 31st. The original price is: input: $0.55/M, output: $2.2/M. The model provider is the Sophnet platform. DeepSeek V3 Fast is a high-TPS, ultra-fast version of DeepSeek V3 0324, featuring full-precision (non-quantized) performance, enhanced code and math capabilities, and faster responses!\n\nDeepSeek V3 0324 is a powerful Mixture-of-Experts (MoE) model with a total parameter count of 671B, activating 37B parameters per token.\nIt adopts Multi-Head Latent Attention (MLA) and the DeepSeekMoE architecture to achieve efficient inference and economical training costs.\nIt innovatively implements a load balancing strategy without auxiliary loss and sets multi-token prediction training targets to enhance performance.\nThe model is pre-trained on 14.8 trillion diverse, high-quality tokens and further optimized through supervised fine-tuning and reinforcement learning stages to fully realize its capabilities.\nComprehensive evaluations show that DeepSeek V3 outperforms other open-source models and rivals leading closed-source models in performance.\nThe entire training process only requires 2.788M H800 GPU hours and remains highly stable, with no irrecoverable loss spikes or rollbacks.","pricing":{"input":0.56,"output":2.24},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":32000,"context_length":32000,"schema_checked":false,"playground_checked":false},{"model_id":"o4-mini","model_name":"O4 Mini","developer_id":12,"desc":"o4-mini is a remarkably smart model for its speed and cost-efficiency. This allows it to support significantly higher usage limits than o3, making it a strong high-volume, high-throughput option for everyone with questions that benefit from reasoning.","pricing":{"cache_read":0.275,"input":1.1,"output":4.4},"types":"llm","features":"thinking,tool,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":100000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4","model_name":"Grok 4","developer_id":9,"desc":"Grok, their latest and greatest flagship model, offers unparalleled performance in natural language, math, and reasoning – the perfect jack of all trades.\nThe current pointing model version is grok-4-0709.","pricing":{"cache_read":0.825,"input":3.3,"output":16.5},"types":"llm","features":"function_calling,structured_outputs,thinking","input_modalities":"text,image","endpoints":"","max_output":64000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-fast-non-reasoning","model_name":"Grok 4 Fast","developer_id":9,"desc":"Grok-4-fast is a cost-effective inference model developed by xAI that delivers cutting-edge performance with excellent token efficiency. The model features a 2 million token context window, advanced Web and X search capabilities, and a unified architecture supporting both \"inference\" and \"non-inference\" modes. Compared to Grok 4, it reduces thinking tokens by an average of 40% and lowers the price by 98% while achieving the same performance.","pricing":{"cache_read":0.05,"input":0.2,"output":0.5},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":30000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4-fast-reasoning","model_name":"Grok 4 Fast (reasoning)","developer_id":9,"desc":"Grok-4-fast is a cost-effective inference model developed by xAI that delivers cutting-edge performance with excellent token efficiency. The model features a 2 million token context window, advanced Web and X search capabilities, and a unified architecture supporting both \"inference\" and \"non-inference\" modes. Compared to Grok 4, it reduces thinking tokens by an average of 40% and lowers the price by 98% while achieving the same performance.","pricing":{"cache_read":0.05,"input":0.2,"output":0.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":30000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2-0711","model_name":"Kimi K2 0711","developer_id":15,"desc":"Kimi-K2 is a MoE architecture foundational model with extremely powerful coding and agent capabilities, featuring a total of 1 trillion parameters and activating 32 billion parameters. In benchmark performance tests across major categories such as general knowledge reasoning, programming, mathematics, and agents, the K2 model outperforms other mainstream open-source models.\nThe Kimi-K2 model supports a context length of 128k tokens.\nIt does not support visual capabilities.","pricing":{"input":0.54,"output":2.16},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":131000,"context_length":131000,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2-instruct","model_name":"Kimi K2 Instruct","developer_id":15,"desc":"Kimi-K2 is a MoE architecture foundational model with extremely powerful coding and agent capabilities, featuring a total of 1 trillion parameters and activating 32 billion parameters. In benchmark performance tests across major categories such as general knowledge reasoning, programming, mathematics, and agents, the K2 model outperforms other mainstream open-source models.\nThe Kimi-K2 model supports a context length of 128k tokens.\nIt does not support visual capabilities.","pricing":{"input":0.54,"output":2.16},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"kimi-k2-turbo-preview","model_name":"Kimi K2 Turbo Preview","developer_id":15,"desc":"The kimi-k2-turbo-preview model is a high-speed version of kimi-k2, with the same model parameters as kimi-k2, but the output speed has been increased from 10 tokens per second to 40 tokens per second.","pricing":{"cache_read":0.3,"input":1.2,"output":4.8},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":262144,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-235b-a22b-instruct","model_name":"Qwen3 VL 235B A22B Instruct","developer_id":13,"desc":"The Qwen3 series open-source models include hybrid models, thinking models, and non-thinking models, with both reasoning capabilities and general abilities reaching industry SOTA levels at the same scale.","pricing":{"input":0.274,"output":1.096},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":33000,"context_length":131000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-235b-a22b-thinking","model_name":"Qwen3 VL 235B A22B Thinking","developer_id":13,"desc":"The Qwen3 series open-source models include hybrid models, thinking models, and non-thinking models, with both reasoning capabilities and general abilities reaching industry SOTA levels at the same scale.","pricing":{"input":0.274,"output":2.74},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":33000,"context_length":131000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-30b-a3b-instruct","model_name":"Qwen3 VL 30B A3B Instruct","developer_id":13,"desc":"The Qwen3-VL series’ second-largest MoE model Instruct version offers fast response speed and supports ultra-long contexts such as long videos and long documents; it features comprehensive upgrades in image/video understanding, spatial perception, and universal recognition abilities; it also provides visual 2DD/3D localization capabilities, making it capable of handling complex real-world tasks.","pricing":{"input":0.1028,"output":0.4112},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-vl-30b-a3b-thinking","model_name":"Qwen3 VL 30B A3B Thinking","developer_id":13,"desc":"The Qwen3-VL series’ second-largest MoE model Thinking version offers fast response speed, stronger multimodal understanding and reasoning, visual agent capabilities, and ultra-long context support for long videos and long documents; it features comprehensive upgrades in image/video understanding, spatial perception, and universal recognition abilities, making it capable of handling complex real-world tasks.","pricing":{"input":0.1028,"output":1.028},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-OCR","model_name":"DeepSeek Ocr","developer_id":7,"desc":"DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical character recognition (OCR) and “contextual optical compression.” The model is designed to explore the limits of compressing contextual information from images, efficiently processing documents and converting them into structured text formats such as Markdown. The model requires an image as input.","pricing":{"input":0.02,"output":0.02},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":8000,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-5.0-thinking-exp","model_name":"ERNIE 5.0 Thinking Exp","developer_id":25,"desc":"ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities.\n\nERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.","pricing":{"cache_read":0.82192,"input":0.82192,"output":3.28768},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":64000,"context_length":119000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1","model_name":"GPT 4.1","developer_id":12,"desc":"The latest flagship multimodal model supports million-token context, with encoding capability (SWE-bench 54.6%) and instruction-following (Scale AI 38.3%) performance significantly surpassing GPT-4o, while reducing costs by 26%, making it suitable for complex tasks. Its automatic caching mechanism offers a 75% cost reduction on cache hits.","pricing":{"cache_read":0.5,"input":2,"output":8},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"aihubmix-router","model_name":"Aihubmix Router","developer_id":12,"desc":"New model routing capability; request aihubmix-router to automatically route models based on question complexity, so everyone no longer needs to manually switch models; in our tests comparing the use of the model router versus only using GPT-4.1, we observed up to 60% cost savings while maintaining similar accuracy.  \nThe context length of the model router depends on the base model used for each prompt. Input size is 200,000, output size is 32,768.  \nCurrently, there are four routing models: gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, o4-mini.  \nPricing: Due to our current billing structure system, requests through aihubmix-router are billed at the price of gpt-4.1-mini regardless of which final model is used; future billing will be based on the actual model invoked.  \nEveryone is welcome to try it out; the interface will return the name of the actual called model.","pricing":{"cache_read":0.1,"input":0.4,"output":1.6},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1-mini","model_name":"GPT 4.1 Mini","developer_id":12,"desc":"Lightweight, high-performance model with million-token context and near-flagship-level encoding and image understanding capabilities, while reducing costs by 83%. It is suitable for rapid development and small to medium-sized applications. The automatic caching mechanism provides a 75% cost reduction on cache hits.","pricing":{"cache_read":0.1,"input":0.4,"output":1.6},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4.1-nano","model_name":"GPT 4.1 Nano","developer_id":12,"desc":"Ultra-lightweight model with million-token context, optimized for speed and low latency, costing only $0.10 per million input tokens. It is suitable for edge computing and real-time interaction. The automatic caching mechanism offers a 75% cost reduction on cache hits.","pricing":{"cache_read":0.025,"input":0.1,"output":0.4},"types":"llm","features":"tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":32768,"context_length":1047576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-05-06","model_name":"Gemini 2.5 Pro Preview 05-06","developer_id":8,"desc":"gemini-2.5-pro latest model","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-03-25","model_name":"Gemini 2.5 Pro Preview 03-25","developer_id":8,"desc":"Supports high concurrency.  \nThe Gemini 2.5 Pro preview version is here, with higher limits for production testing.  \nGoogle's latest and most powerful model;","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-05-06-search","model_name":"Gemini 2.5 Pro Preview 05-06 Search","developer_id":8,"desc":"Integrated with Google's official search function.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm,search","features":"thinking,web","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-03-25-search","model_name":"Gemini 2.5 Pro Preview 03-25 Search","developer_id":8,"desc":"Integrated with Google's official search function.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm,search","features":"thinking,web,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-max-preview","model_name":"Qwen3 Max Preview","developer_id":13,"desc":"Qwen3-Max-Preview is the latest preview model in the Qwen3 series. This version is functionally equivalent to Qwen3-Max-Thinking — simply set extra_body={\"enable_thinking\": True} to enable the thinking mode. Compared to the Qwen2.5 series, it delivers significant improvements in overall general capabilities, including English–Chinese text understanding, complex instruction following, open-ended reasoning, multilingual processing, and tool-use proficiency. The model also exhibits fewer hallucinations and stronger overall reliability.","pricing":{"cache_read":0.1692,"input":0.846,"output":3.384},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-max","model_name":"Qwen3 Max","developer_id":13,"desc":"The Tongyi Qianwen 3 series Max model has undergone special upgrades in intelligent agent programming and tool invocation compared to the preview version. The officially released model this time reaches SOTA level in the field and is adapted to more complex intelligent agent scenarios.","pricing":{"cache_read":0.09016,"cache_write":0.5635,"input":0.4508,"output":1.8032},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":65536,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-next-80b-a3b-instruct","model_name":"Qwen3 Next 80B A3B Instruct","developer_id":13,"desc":"Qwen3-Next-80B-A3B-Instruct is an instruction-tuned model in the Qwen3-Next series, optimized for delivering fast, stable, and direct final answers without showing its reasoning steps (\"thinking traces\").\n\nUnlike chain-of-thought models, it focuses on generating consistent, instruction-following outputs, making it ideal for production environments. It excels at complex tasks like reasoning and coding while maintaining high throughput and stability, especially with ultra-long inputs and multi-turn dialogues.\n\nEngineered for efficiency, its performance rivals larger Qwen3 systems, making it perfectly suited for RAG, tool use, and agentic workflows where deterministic results are critical.","pricing":{"input":0.138,"output":0.552},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-next-80b-a3b-thinking","model_name":"Qwen3 Next 80B A3B Thinking","developer_id":13,"desc":"Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that excels by outputting structured 'thinking' traces (Chain-of-Thought) by default.\n\nDesigned for hard, multi-step problems, it is ideal for tasks like math proofs, code synthesis, logic puzzles, and agentic planning. Compared to other Qwen3 variants, it offers greater stability during long reasoning chains and is tuned to follow complex instructions without getting repetitive or off-task.\n\nThis model is perfectly suited for agent frameworks, tool use (function calling), and benchmarks where a step-by-step breakdown is required. It leverages throughput-oriented techniques for fast generation of detailed, procedural outputs.","pricing":{"input":0.142,"output":1.42},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-235b-a22b-instruct-2507","model_name":"Qwen3 235B A22B Instruct 2507","developer_id":13,"desc":"Qwen3-235B-A22B-Instruct-2507","pricing":{"input":0.28,"output":1.12},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":262144,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-235b-a22b-thinking-2507","model_name":"Qwen3 235B A22B Thinking 2507","developer_id":13,"desc":"The open-source thinking model based on Qwen3 has significantly improved in logical ability, general capability, knowledge enhancement, and creative ability compared to the previous version (Tongyi Qianwen 3-235B-A22B). It is suitable for high-difficulty and strong reasoning scenarios.","pricing":{"input":0.28,"output":2.8},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":262144,"context_length":262144,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-30b-a3b-instruct","model_name":"Qwen3 Coder 30B A3B Instruct","developer_id":13,"desc":"The code generation model based on Qwen3 has powerful Coding Agent capabilities, achieving state-of-the-art performance compared to open-source models.The model adopts tiered pricing.","pricing":{"cache_read":0.2,"input":0.2,"output":0.8},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":262000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-480b-a35b-instruct","model_name":"Qwen3 Coder 480B A35B Instruct","developer_id":13,"desc":"The code generation model based on Qwen3 has powerful Coding Agent capabilities, achieving state-of-the-art performance compared to open-source models.The model adopts tiered pricing.","pricing":{"cache_read":0.82,"input":0.82,"output":3.28},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":262000,"context_length":262000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-235b-a22b","model_name":"Qwen3 235B A22B","developer_id":13,"desc":"Qwen3-235B-A22B is a massive 235B parameter Mixture-of-Experts (MoE) model that operates with the efficiency of a 22B model. Its standout feature is the ability to seamlessly switch between a \"thinking\" mode for complex reasoning and a \"non-thinking\" mode for fast conversation, offering both world-class power and practical speed.","pricing":{"input":0.28,"output":1.12},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":131100,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-flash","model_name":"Qwen3 Coder Flash","developer_id":13,"desc":"Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling and environment interaction, combining coding proficiency with versatile general-purpose abilities.","pricing":{"cache_read":0.136,"input":0.136,"output":0.544},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":65536,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-plus","model_name":"Qwen3 Coder Plus","developer_id":13,"desc":"The code generation model based on Qwen3 has powerful Coding Agent capabilities, excels in tool invocation and environment interaction, and can achieve autonomous programming with outstanding coding abilities while also possessing general capabilities.The model adopts tiered pricing.","pricing":{"cache_read":0.108,"input":0.54,"output":2.16},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-coder-plus-2025-07-22","model_name":"Qwen3 Coder Plus 2025 07-22","developer_id":13,"desc":"The code generation model based on Qwen3 has powerful Coding Agent capabilities, excels in tool invocation and environment interaction, and can achieve autonomous programming with outstanding coding abilities while also possessing general capabilities.The model adopts tiered pricing.","pricing":{"cache_read":0.54,"input":0.54,"output":2.16},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":65536,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-V3","model_name":"DeepSeek V3","developer_id":7,"desc":"It has been automatically upgraded to the latest released version, 250324.\nAutomatically upgraded to the latest released version 250324.","pricing":{"input":0.272,"output":1.088},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":1638000,"context_length":1638000,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-06-05-search","model_name":"Gemini 2.5 Pro Preview 06-05 Search","developer_id":8,"desc":"Integrated with Google's official search function.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm,search","features":"thinking,web,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-5.0-thinking-preview","model_name":"ERNIE 5.0 Thinking Preview","developer_id":25,"desc":"The new generation Wenxin model, Wenxin 5.0, is a native full-modal large model that adopts native full-modal unified modeling technology, jointly modeling text, images, audio, and video, possessing comprehensive full-modal capabilities. Wenxin 5.0's basic abilities are comprehensively upgraded, performing excellently on benchmark test sets, especially in multimodal understanding, instruction compliance, creative writing, factual accuracy, intelligent agent planning, and tool application.","pricing":{"cache_read":0.822,"input":0.822,"output":3.288},"types":"llm","features":"thinking,structured_outputs,function_calling","input_modalities":"text","endpoints":"","max_output":64000,"context_length":183000,"schema_checked":false,"playground_checked":false},{"model_id":"inclusionAI/Ling-1T","model_name":"Ling 1t","developer_id":29,"desc":"Ling-1T is the first flagship non-thinking model in the “Ling 2.0” series, featuring 1 trillion total parameters and approximately 50 billion active parameters per token. Built on the Ling 2.0 architecture, Ling-1T is designed to push the limits of efficient inference and scalable cognition. Ling-1T-base was pretrained on over 20 trillion high-quality, reasoning-intensive tokens, supports up to a 128K context length, and incorporates an Evolutionary Chain of Thought (Evo-CoT) process during mid-stage and post-stage training. This training regimen greatly enhances the model’s efficiency and depth of reasoning, enabling Ling-1T to achieve top performance across multiple complex reasoning benchmarks, balancing accuracy and efficiency.","pricing":{"input":0.548,"output":2.192},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"inclusionAI/Ring-1T","model_name":"Ring 1t","developer_id":29,"desc":"Ring-1T is an open-source idea model with a trillion parameters released by the Bailing team. It is based on the Ling 2.0 architecture and the Ling-1T-base foundational model for training, with a total parameter count of 1 trillion, an active parameter count of 50 billion, and supports up to a 128K context window. The model is trained via large-scale verifiable reward reinforcement learning (RLVR), combined with the self-developed Icepop reinforcement learning stabilization method and the efficient ASystem reinforcement learning system, significantly improving the model’s deep reasoning and natural language reasoning capabilities. Ring-1T achieves leading performance among open-source models on high-difficulty reasoning benchmarks such as mathematics competitions (e.g., IMO 2025), code generation (e.g., ICPC World Finals 2025), and logical reasoning.","pricing":{"input":0.548,"output":2.192},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"inclusionAI/Ling-flash-2.0","model_name":"Ling Flash 2.0","developer_id":29,"desc":"Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion parameters, of which 6.1 billion are activated per token (4.8 billion non-embedding). As part of the Ling 2.0 architecture series, it is designed as a lightweight yet powerful Mixture-of-Experts (MoE) model. It aims to deliver performance comparable to or even exceeding that of 40B-level dense models and other larger MoE models, but with a significantly smaller active parameter count. The model represents a strategy focused on achieving high performance and efficiency through extreme architectural design and training methods.","pricing":{"input":0.136,"output":0.544},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"inclusionAI/Ling-mini-2.0","model_name":"Ling Mini 2.0","developer_id":29,"desc":"Ling-mini-2.0 is a small-sized, high-performance large language model based on the MoE architecture. It has a total of 16 billion parameters, but only activates 1.4 billion parameters per token (non-embedding 789 million), achieving extremely high generation speed. Thanks to the efficient MoE design and large-scale high-quality training data, despite activating only 1.4 billion parameters, Ling-mini-2.0 still demonstrates top-tier performance on downstream tasks comparable to dense LLMs under 10 billion parameters and even larger-scale MoE models.","pricing":{"input":0.068,"output":0.272},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"inclusionAI/Ring-flash-2.0","model_name":"Ring Flash 2.0","developer_id":29,"desc":"Ring-flash-2.0 is a high-performance thinking model deeply optimized based on the Ling-flash-2.0-base. It uses a mixture-of-experts (MoE) architecture with a total of 100 billion parameters, but only activates 6.1 billion parameters per inference. The model employs the original Icepop algorithm to solve the instability issues of large MoE models during reinforcement learning (RL) training, enabling its complex reasoning capabilities to continuously improve over long training cycles. Ring-flash-2.0 has achieved significant breakthroughs on multiple high-difficulty benchmarks, including mathematics competitions, code generation, and logical reasoning. Its performance not only surpasses top dense models under 40 billion parameters but also rivals larger open-source MoE models and closed-source high-performance thinking models. Although the model focuses on complex reasoning, it also performs exceptionally well on creative writing tasks. Furthermore, thanks to its efficient architecture, Ring-flash-2.0 delivers high performance with low-latency inference, significantly reducing deployment costs in high-concurrency scenarios.","pricing":{"input":0.136,"output":0.544},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"jina-deepsearch-v1","model_name":"Jina Deepsearch V1","developer_id":22,"desc":"DeepSearch combines search, reading, and reasoning capabilities to pursue the best possible answer. It's fully compatible with OpenAI's Chat API format—just replace api.openai.com with aihubmix.com to get started.  \nThe stream will return the thinking process.","pricing":{"input":0.05,"output":0.05},"types":"llm,search","features":"thinking,web,deepsearch","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"llama-4-maverick","model_name":"Llama 4 Maverick","developer_id":11,"desc":"Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring 400B total parameters and 128 experts, while activating an efficient 17B parameters per inference. Engineered for peak performance, it excels at advanced multimodal tasks.\n\nMaverick natively supports text and image input, producing multilingual text and code. With a 1-million-token context window and instruction tuning, it is optimized for complex image reasoning and general-purpose assistant-like interactions.\n\nReleased under the Llama 4 Community License, Maverick is ideal for research and commercial applications demanding state-of-the-art multimodal understanding and high throughput.","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"llama-4-scout","model_name":"Llama 4 Scout","developer_id":11,"desc":"Llama 4 Scout is a highly efficient Mixture-of-Experts (MoE) model from Meta, activating 17B out of 109B total parameters per inference. It natively supports multimodal input (text and image) and multilingual output (text and code) across 12 languages.\n\nDesigned for assistant-style interaction and visual reasoning, Scout features a massive 10-million-token context window. It is instruction-tuned for tasks like multilingual chat and image understanding and is released under the Llama 4 Community License for local or commercial deployment.","pricing":{"input":0.2,"output":0.2},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":131000,"context_length":131000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-mt-plus","model_name":"Qwen Mt Plus","developer_id":13,"desc":"Based on the comprehensive upgrade of Qwen3, this flagship translation large model supports bidirectional translation across 92 languages. It offers fully enhanced model performance and translation quality, along with more stable terminology customization, format fidelity, and domain-prompting capabilities, making translations more accurate and natural.","pricing":{"input":0.492,"output":1.476},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":8000,"context_length":16000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-mt-turbo","model_name":"Qwen Mt Turbo","developer_id":13,"desc":"Based on the comprehensive upgrade of Qwen3, this flagship translation large model supports bidirectional translation across 92 languages. It offers fully enhanced model performance and translation quality, along with more stable terminology customization, format fidelity, and domain-prompting capabilities, making translations more accurate and natural.","pricing":{"input":0.192,"output":0.534912},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":8000,"context_length":16000,"schema_checked":false,"playground_checked":false},{"model_id":"codex-mini-latest","model_name":"Codex Mini","developer_id":12,"desc":"Only supports v1/responses API calls.https://docs.aihubmix.com/en/api/Responses-API\ncodex-mini-latest is a fine-tuned version of o4-mini specifically for use in Codex CLI. For direct use in the API, we recommend starting with gpt-4.1.","pricing":{"cache_read":0.375,"input":1.5,"output":6},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-4.5-turbo-latest","model_name":"ERNIE 4.5 Turbo","developer_id":25,"desc":"Wenxin 4.5 Turbo also has significant improvements in hallucination reduction, logical reasoning, and coding capabilities. Compared to Wenxin 4.5, it is faster and more affordable.","pricing":{"input":0.11,"output":0.44},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":12000,"context_length":135000,"schema_checked":false,"playground_checked":false},{"model_id":"DeepSeek-R1","model_name":"DeepSeek R1","developer_id":7,"desc":"DeepSeek R1 is a new open-source model with performance on par with OpenAI's o1 and features fully open reasoning tokens. It is a 671B-parameter Mixture-of-Experts (MoE) model that activates 37B parameters during inference.","pricing":{"input":0.4,"output":2},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":1638000,"context_length":1638000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-search-preview","model_name":"GPT 4o Search Preview","developer_id":12,"desc":"Using the Chat Completions API, you can directly access the fine-tuned models and tool used by Search in ChatGPT.\n\nWhen using Chat Completions, the model always retrieves information from the web before responding to your query. To use web_search_preview as a tool that models like gpt-4o and gpt-4o-mini invoke only when necessary, switch to using the Responses API.\n\nCurrently, you need to use one of these models to use web search in Chat Completions:\n\ngpt-4o-search-preview\ngpt-4o-mini-search-preview\nWeb search parameter example\nimport OpenAI from \"openai\";\nconst client = new OpenAI();\n\nconst completion = await client.chat.completions.create({\n    model: \"gpt-4o-search-preview\",\n    web_search_options: {},\n    messages: [{\n        \"role\": \"user\",\n        \"content\": \"What was a positive news story from today?\"\n    }],\n});\n\nconsole.log(completion.choices[0].message.content);\nOutput and citations\nThe API response item in the choices array will include:\n\nmessage.content with the text result from the model, inclusive of any inline citations\nannotations with a list of cited URLs\nBy default, the model's response will include inline citations for URLs found in the web search results. In addition to this, the url_citation annotation object will contain the URL and title of the cited source, as well as the start and end index characters in the model's response where those sources were used.","pricing":{"cache_read":1.25,"input":2.5,"output":10},"types":"llm,search","features":"web,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-mini-search-preview","model_name":"GPT 4o Mini Search Preview","developer_id":12,"desc":"Using the Chat Completions API, you can directly access the fine-tuned models and tool used by Search in ChatGPT.\n\nWhen using Chat Completions, the model always retrieves information from the web before responding to your query. To use web_search_preview as a tool that models like gpt-4o and gpt-4o-mini invoke only when necessary, switch to using the Responses API.\n\nCurrently, you need to use one of these models to use web search in Chat Completions:\n\ngpt-4o-search-preview\ngpt-4o-mini-search-preview\nWeb search parameter example\nimport OpenAI from \"openai\";\nconst client = new OpenAI();\n\nconst completion = await client.chat.completions.create({\n    model: \"gpt-4o-search-preview\",\n    web_search_options: {},\n    messages: [{\n        \"role\": \"user\",\n        \"content\": \"What was a positive news story from today?\"\n    }],\n});\n\nconsole.log(completion.choices[0].message.content);\nOutput and citations\nThe API response item in the choices array will include:\n\nmessage.content with the text result from the model, inclusive of any inline citations\nannotations with a list of cited URLs\nBy default, the model's response will include inline citations for URLs found in the web search results. In addition to this, the url_citation annotation object will contain the URL and title of the cited source, as well as the start and end index characters in the model's response where those sources were used.","pricing":{"cache_read":0.075,"input":0.15,"output":0.6},"types":"llm,search","features":"web,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-3-7-sonnet","model_name":"Claude 3.7 Sonnet","developer_id":2,"desc":"Support for the thinking parameter through the original Claude SDK.","pricing":{"input":3.3,"output":16.5},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":128000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-4.5","model_name":"ERNIE 4.5","developer_id":25,"desc":"Wenxin Large Model 4.5 is a next-generation native multimodal foundational model independently developed by Baidu. It achieves collaborative optimization through joint modeling of multiple modalities, demonstrating excellent multimodal understanding capabilities; it possesses more advanced language abilities, with comprehensive improvements in comprehension, generation, logic, and memory, as well as significant enhancements in hallucination reduction, logical reasoning, and coding capabilities.ERNIE-4.5-21B-A3B is an aligned open-source model with a MoE structure, having a total of 21 billion parameters and 3 billion activated parameters.","pricing":{"input":0.068,"output":0.272},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":64000,"context_length":160000,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-4.5-turbo-vl","model_name":"ERNIE 4.5 Turbo VL","developer_id":25,"desc":"The new version of the Wenxin Yiyan large model significantly improves capabilities in image understanding, creation, translation, and coding. It supports a context length of up to 32K tokens for the first time, with a notable reduction in the latency of the first token.","pricing":{"input":0.4,"output":1.2},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16000,"context_length":139000,"schema_checked":false,"playground_checked":false},{"model_id":"mimo-v2-flash-free","model_name":"MiMo V2 Flash (free)","developer_id":31,"desc":"MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It adopts a MoE architecture with 309B total parameters and 15B active parameters per inference, balancing performance and efficiency. The model features a hybrid attention architecture, supports a hybrid-thinking toggle, and offers a 256K context window, enabling strong capabilities in complex reasoning, code generation, and agent-based scenarios. On SWE-bench Verified and SWE-bench Multilingual, MiMo-V2-Flash ranks #1 among open-source models globally, delivering performance comparable to Claude Sonnet 4.5 while costing only about 3.5% as much.","pricing":{"cache_read":0,"input":0,"output":0},"types":"llm","features":"web","input_modalities":"text","endpoints":"","max_output":256000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"o3-mini","model_name":"O3 Mini","developer_id":12,"desc":"OpenAI's latest fast inference model excels at STEAM tasks and offers exceptional cost-effectiveness. Official support for cache hits reduces input prices by half.","pricing":{"cache_read":0.55,"input":1.1,"output":4.4},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":100000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-1-6","model_name":"Doubao Seed 1.6","developer_id":4,"desc":"Doubao-Seed-1.6 is a brand new multimodal deep reasoning model that supports four types of reasoning effort: minimal, low, medium, and high. It offers stronger model performance, serving complex tasks and challenging scenarios. It supports a 256k context window, with output length up to a maximum of 32k tokens.","pricing":{"cache_read":0.036,"input":0.18,"output":1.8},"types":"llm","features":"thinking，tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-1-6-flash","model_name":"Doubao Seed 1.6 Flash","developer_id":4,"desc":"Doubao-Seed-1.6-flash is an extremely fast multimodal deep thinking model, with TPOT requiring only 10ms. It supports both text and visual understanding, with its text comprehension skills surpassing the previous generation lite model and its visual understanding on par with competitor's pro series models. It supports a 256k context window and an output length of up to 16k tokens.","pricing":{"cache_read":0.0088,"input":0.044,"output":0.44},"types":"llm","features":"thinking，tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":33000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-1-6-lite","model_name":"Doubao Seed 1.6 Lite","developer_id":4,"desc":"Doubao-Seed-1.6-lite is a brand new multimodal deep reasoning model that supports adjustable reasoning effort, with four modes: Minimal, Low, Medium, and High. It offers better cost performance, making it the best choice for common tasks, with a context window of up to 256k.","pricing":{"cache_read":0.0164,"input":0.082,"output":0.656},"types":"llm","features":"thinking，tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"doubao-seed-1-6-thinking","model_name":"Doubao Seed 1.6 Thinking","developer_id":4,"desc":"The Doubao-Seed-1.6-thinking model has significantly enhanced reasoning capabilities. Compared with Doubao-1.5-thinking-pro, it has further improvements in fundamental abilities such as coding, mathematics, and logical reasoning, and now also supports visual understanding. It supports a 256k context window, with output length supporting up to 16k tokens.","pricing":{"cache_read":0.036,"input":0.18,"output":1.8},"types":"llm","features":"thinking，tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-oss-120b","model_name":"gpt-oss-120b","developer_id":12,"desc":"gpt-oss-120b is a 117B-parameter open-weight Mixture-of-Experts (MoE) language model from OpenAI, designed for high-reasoning, agentic, and general-purpose production use cases. Activating just 5.1B parameters per pass, it is optimized to run on a single H100 GPU with native MXFP4 quantization. The model features configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.","pricing":{"input":0.18,"output":0.9},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":32768,"context_length":131072,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-preview-06-05","model_name":"Gemini 2.5 Pro Preview 06-05","developer_id":8,"desc":"Google’s latest multimodal flagship model, combining exceptional coding and reasoning capabilities. Its massive 1 million token context window (soon to expand to 2 million) places it at the top of the WebDevArena and LMArena leaderboards. It is particularly well-suited for developing aesthetically pleasing and highly functional interactive web applications, code transformation, and complex workflows. The newly introduced \"reasoning budget\" feature cleverly balances cost and performance, while optimized tool calls and response styles further enhance development efficiency, making it the ideal choice for rapid prototyping and advanced coding.","pricing":{"cache_read":0.125,"input":1.25,"output":10},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":65536,"context_length":1048576,"schema_checked":false,"playground_checked":false},{"model_id":"Qwen/Qwen2.5-VL-72B-Instruct","model_name":"Qwen2.5 VL 72B Instruct","developer_id":13,"desc":"Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong visual understanding and reasoning capabilities. It can recognize objects, analyze text and charts, understand key events in long videos, and accurately locate targets within images. The model supports structured output, making it suitable for data such as invoices and forms, and performs excellently in multiple benchmark tests.","pricing":{"cache_read":0,"input":0.5,"output":0.5},"types":"llm","features":"","input_modalities":"text,image,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"o1","model_name":"O1","developer_id":12,"desc":"OpenAI's most powerful O-series model supports official cache hits that halve the input cost.","pricing":{"cache_read":7.5,"input":15,"output":60},"types":"llm","features":"thinking","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"o1-pro","model_name":"O1 Pro","developer_id":12,"desc":"The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide consistently better answers.","pricing":{"cache_read":170,"input":170,"output":680},"types":"llm","features":"thinking","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ByteDance-Seed/Seed-OSS-36B-Instruct","model_name":"Seed Oss 36B Instruct","developer_id":4,"desc":"Seed-OSS is a series of open-source large language models developed by ByteDance's Seed team, designed specifically for powerful long-context processing, reasoning, agents, and general capabilities. Among this series, Seed-OSS-36B-Instruct is an instruction-tuned model with 36 billion parameters that natively supports ultra-long context lengths, enabling it to process massive documents or complex codebases in a single pass. This model is specially optimized for reasoning, code generation, and agent tasks (such as tool usage), while maintaining balanced and excellent general capabilities. A notable feature of this model is the \"Thinking Budget\" functionality, which allows users to flexibly adjust the inference length as needed, thereby effectively improving inference efficiency in practical applications.","pricing":{"input":0.2,"output":0.534},"types":"llm","features":"thinking，tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":256000,"schema_checked":false,"playground_checked":false},{"model_id":"cc-minimax-m2","model_name":"CC MiniMax M2","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"deepseek-r1-distill-llama-70b","model_name":"DeepSeek R1 Distill Llama 70B","developer_id":7,"desc":"Provided by Groq, the DeepSeek-R1-Distill model is fine-tuned based on an open-source model, using samples generated by DeepSeek-R1. We have made slight modifications to their configurations and tokenizers. Please use our settings to run these models.","pricing":{"input":0.8,"output":1.6},"types":"llm","features":"thinking","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"o1-preview","model_name":"O1 Preview","developer_id":12,"desc":"The latest and most powerful inference model from OpenAI; AiHubMix uses both OpenAI and Microsoft Azure OpenAI channels simultaneously to achieve high-concurrency load balancing.","pricing":{"cache_read":7.5,"input":15,"output":60},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-2024-11-20","model_name":"GPT 4o 2024 11-20","developer_id":12,"desc":"The latest version of the GPT-4o model; it is recommended to use this version, as it is currently smarter than the regular 4o.","pricing":{"cache_read":1.25,"input":2.5,"output":10},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o","model_name":"GPT 4o","developer_id":12,"desc":"GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more natural human–computer interaction. It can accept any combination of text, audio, image, and video as input, and generate multimodal outputs including text, audio, and images. With audio response latency as low as 232 milliseconds on average around 320 milliseconds, it approaches real human conversational speed. The model delivers strong performance in English text and code, significantly improved multilingual understanding, and outstanding capabilities in visual and audio perception, while offering faster API performance and substantially reduced cost for real-time and complex multimodal applications.","pricing":{"cache_read":1.25,"input":2.5,"output":10},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-mini","model_name":"GPT 4o Mini","developer_id":12,"desc":"The lightweight version of GPT-4o, which is affordable and fast, suitable for handling simple tasks; our site supports the official automatic caching for this model, and charges for cache hits will be automatically halved.","pricing":{"cache_read":0.075,"input":0.15,"output":0.6},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":16384,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"ERNIE-X1.1-Preview","model_name":"ERNIE X1.1 Preview","developer_id":25,"desc":"The Wenxin large model X1.1 has made significant improvements in question answering, tool invocation, intelligent agents, instruction following, logical reasoning, mathematics, and coding tasks, with notable enhancements in factual accuracy. The context length has been extended to 64K tokens, supporting longer inputs and dialogue history, which improves the coherence of long-chain reasoning while maintaining response speed.","pricing":{"input":0.136,"output":0.544},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":64000,"context_length":119000,"schema_checked":false,"playground_checked":false},{"model_id":"Qwen/QwQ-32B","model_name":"QwQ 32B","developer_id":13,"desc":"Silicon-based flow provision","pricing":{"input":0.14,"output":0.56},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"minimax-m2","model_name":"MiniMax M2","developer_id":18,"desc":"MiniMax-M2 redefines efficiency for intelligent agents. It is a compact, fast, and cost-effective MoE model with a total of 230 billion parameters and 10 billion active parameters, designed for top performance in coding and intelligent agent tasks while maintaining strong general intelligence. With only 10 billion active parameters, MiniMax-M2 delivers the complex end-to-end tool usage performance expected from today's leading models, but in a more streamlined form factor, making deployment and scaling easier than ever before.","pricing":{"input":0.288,"output":1.152},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":192000,"context_length":204800,"schema_checked":false,"playground_checked":false},{"model_id":"kat-dev","model_name":"Kat Dev","developer_id":13,"desc":"KAT-Dev (32B) is an open-source 32B parameter model specifically designed for software engineering tasks. It achieved a 62.4% resolution rate on the SWE-Bench Verified benchmark, ranking fifth among all open-source models of various scales. The model is optimized through multiple stages, including intermediate training, supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT), as well as large-scale agent reinforcement learning (RL). Based on Qwen3-32B, its training process lays the foundation for subsequent fine-tuning and reinforcement learning stages by enhancing fundamental abilities such as tool usage, multi-turn interaction, and instruction following. During the fine-tuning phase, the model not only learns eight carefully curated task types and programming scenarios but also innovatively introduces a reinforcement fine-tuning (RFT) stage guided by human engineer-annotated “teacher trajectories.” The final agent reinforcement learning phase addresses scalability challenges through multi-level prefix caching, entropy-based trajectory pruning, and efficient architecture.","pricing":{"input":0.137,"output":0.548},"types":"llm","features":"tools","input_modalities":"text","endpoints":"","max_output":0,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"nvidia-nemotron-3-super-120b-a12b","model_name":"Nvidia Nemotron 3 Super 120B A12B","developer_id":17,"desc":"An open-source, efficient hybrid Mamba-Transformer MoE model that supports a context length of one million tokens and excels at agent reasoning, programming, planning, and tool invocation.","pricing":{"cache_read":0.0275,"input":0.11,"output":0.55},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text","endpoints":"","max_output":0,"context_length":1000000,"schema_checked":false,"playground_checked":false},{"model_id":"qwen2.5-vl-72b-instruct","model_name":"Qwen2.5 VL 72B Instruct","developer_id":13,"desc":"Strong capability in Chinese domain recognition, comparable to ChatGPT-4.0.","pricing":{"input":2.4,"output":7.2},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"Qwen/Qwen2.5-VL-32B-Instruct","model_name":"Qwen2.5 VL 32B Instruct","developer_id":13,"desc":"Qwen2.5-VL-32B-Instruct is an advanced multimodal model from the Tongyi Qianwen team that can recognize objects, analyze text and graphics in images, operate tools, locate objects in images, and generate structured outputs. Through reinforcement learning, it has improved mathematics and problem-solving capabilities, with a more concise and natural response style.","pricing":{"input":0.24,"output":0.24},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"baidu/ERNIE-4.5-300B-A47B","model_name":"ERNIE 4.5 300B A47B","developer_id":25,"desc":"ERNIE-4.5-300B-A47B is a large language model developed by Baidu based on a Mixture of Experts (MoE) architecture. The model has a total of 300 billion parameters, but only activates 47 billion parameters per token during inference, which balances strong performance with computational efficiency. As one of the core models in the ERNIE 4.5 series, it demonstrates outstanding capabilities in tasks such as text understanding, generation, reasoning, and programming. The model employs an innovative multimodal heterogeneous MoE pretraining approach, leveraging joint training of textual and visual modalities to effectively enhance the model’s overall abilities, particularly excelling in instruction following and world knowledge memorization. Baidu has open-sourced this model along with other models in the series, aiming to promote the research and application of AI technology.","pricing":{"cache_read":0,"input":0.32,"output":1.28},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-4.5-0.3b","model_name":"ERNIE 4.5 0.3b","developer_id":25,"desc":"Wenxin Large Model 4.5 is a next-generation native multimodal foundational large model independently developed by Baidu. It achieves collaborative optimization through joint modeling of multiple modalities, demonstrating excellent multimodal understanding capabilities. The model possesses enhanced language abilities, with comprehensive improvements in understanding, generation, reasoning, and memory. It significantly reduces hallucinations and shows notable advancements in logical reasoning and coding skills.","pricing":{"input":0.0136,"output":0.0544},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-4.5-turbo-128k-preview","model_name":"ERNIE 4.5 Turbo 128K Preview","developer_id":25,"desc":"Wenxin 4.5 Turbo also shows significant enhancements in reducing hallucinations, logical reasoning, and coding capabilities. Compared to Wenxin 4.5, it is faster and more cost-effective.","pricing":{"input":0.108,"output":0.432},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ernie-x1-turbo","model_name":"ERNIE X1 Turbo","developer_id":25,"desc":"Wenxin Large Model X1 possesses enhanced abilities in understanding, planning, reflection, and evolution. As a more comprehensive deep-thinking model, Wenxin X1 combines accuracy, creativity, and literary elegance, excelling particularly in Chinese knowledge Q\u0026A, literary creation, document writing, daily conversations, logical reasoning, complex calculations, and tool invocation.","pricing":{"input":0.136,"output":0.544},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":28000,"context_length":50500,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-zh","model_name":"GPT 4o Zh","developer_id":12,"desc":"","pricing":{"input":2.5,"output":10},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-plus-latest","model_name":"Qwen Plus","developer_id":13,"desc":"The Qwen series models with balanced capabilities have inference performance and speed between Qwen-Max and Qwen-Turbo, making them suitable for moderately complex tasks. This model is a dynamically updated version, and updates will not be announced in advance. The current version is qwen-plus-2025-04-28.The model adopts tiered pricing.","pricing":{"cache_read":0.02252,"cache_write":0.14075,"input":0.1126,"output":1.126},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-turbo-latest","model_name":"Qwen Turbo","developer_id":13,"desc":"The Qwen series model with the fastest speed and lowest cost, suitable for simple tasks. This model is a dynamically updated version, and updates will not be announced in advance. The model's overall Chinese and English abilities have been significantly improved, human preference alignment has been greatly enhanced, inference capability and complex instruction understanding have been substantially strengthened, performance on difficult tasks is better, and mathematics and coding skills have been significantly improved. The current version is qwen-turbo-2025-04-28.","pricing":{"cache_read":0.0092,"input":0.046,"output":0.092},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"AiHubmix-Phi-4-mini-reasoning","model_name":"Aihubmix Phi 4 Mini (reasoning)","developer_id":3,"desc":"Phi-4-mini-reasoning is a lightweight open model designed for advanced mathematical reasoning and logic-intensive problem-solving. It is particularly well-suited for tasks such as formal proofs, symbolic computation, and solving multi-step word problems. With its efficient architecture, the model balances high-quality reasoning performance with cost-effective deployment, making it ideal for educational applications, embedded tutoring, and lightweight edge or mobile systems.\n\nPhi-4-mini-reasoning supports a 128K token context length, enabling it to process and reason over long mathematical problems and proofs. Built on synthetic and high-quality math datasets, the model leverages advanced fine-tuning techniques such as supervised fine-tuning and preference modeling to enhance reasoning capabilities. Its training incorporates safety and alignment protocols, ensuring robust and reliable performance across supported use cases.","pricing":{"input":0.12,"output":0.12},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":4000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"aihub-Phi-4-multimodal-instruct","model_name":"Aihub Phi 4 Multimodal Instruct","developer_id":3,"desc":"Microsoft's latest model","pricing":{"input":0.12,"output":0.48},"types":"llm","features":"","input_modalities":"text,image,audio","endpoints":"","max_output":4000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"aihub-Phi-4-mini-instruct","model_name":"Aihub Phi 4 Mini Instruct","developer_id":3,"desc":"Microsoft's latest model","pricing":{"input":0.12,"output":0.48},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":4000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"aihub-Phi-4","model_name":"Aihub Phi 4","developer_id":3,"desc":"Phi-4 is a state-of-the-art open model based on a combination of synthetic datasets, curated public domain website data, and acquired academic books and QA datasets. The approach aims to ensure that small, efficient models are trained using data focused on high quality and advanced reasoning.","pricing":{"input":0.12,"output":0.48},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":16400,"context_length":16400,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-14b","model_name":"Qwen3 14B","developer_id":13,"desc":"Achieves effective integration of thinking and non-thinking modes, enabling mode switching during conversations. Its reasoning ability reaches state-of-the-art (SOTA) levels among models of the same scale, and its general capability significantly surpasses Qwen2.5-14B.","pricing":{"cache_read":0,"input":0.16,"output":1.6},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-8b","model_name":"Qwen3 8B","developer_id":13,"desc":"Achieves effective integration of thinking and non-thinking modes, enabling mode switching during conversations. Its reasoning ability reaches state-of-the-art (SOTA) levels among models of the same scale, and its general capability significantly surpasses Qwen2.5-7B.","pricing":{"cache_read":0,"input":0.08,"output":0.8},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-4b","model_name":"Qwen3 4B","developer_id":13,"desc":"Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning ability reaches state-of-the-art (SOTA) levels among models of the same scale, with significantly enhanced human preference alignment. There are notable improvements in creative writing, role-playing, multi-turn dialogue, and instruction following, resulting in a noticeably better user experience.","pricing":{"cache_read":0,"input":0.046,"output":0.46},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-1.7b","model_name":"Qwen3 1.7b","developer_id":13,"desc":"Effectively integrates thinking and non-thinking modes, allowing mode switching during conversations. Its general capabilities significantly surpass those of the Qwen2.5 small-scale series, with greatly enhanced human preference alignment. There are notable improvements in creative writing, role-playing, multi-turn dialogue, and instruction following, resulting in a significantly better expected user experience.","pricing":{"cache_read":0,"input":0.046,"output":0.46},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen3-0.6b","model_name":"Qwen3 0.6b","developer_id":13,"desc":"Effectively integrates thinking and non-thinking modes, allowing mode switching during conversations. Its general capabilities significantly surpass those of the Qwen2.5 small-scale series.","pricing":{"cache_read":0,"input":0.046,"output":0.46},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-turbo-2025-04-28","model_name":"Qwen Turbo 2025 04-28","developer_id":13,"desc":"The Qwen3 series Turbo model effectively integrates thinking and non-thinking modes, allowing seamless switching between modes during conversations. With a smaller parameter size, its reasoning ability rivals that of QwQ-32B, and its general capabilities significantly surpass those of Qwen2.5-Turbo, reaching state-of-the-art (SOTA) levels among models of the same scale. This version is a snapshot model as of April 28, 2025.","pricing":{"cache_read":0,"input":0.046,"output":0.092},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command-a-03-2025","model_name":"Command A 03 2025","developer_id":6,"desc":"Command A is Cohere most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.","pricing":{"cache_read":0,"input":2.5,"output":10},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-plus-2025-04-28","model_name":"Qwen Plus 2025 04-28","developer_id":13,"desc":"The Qwen3 series Plus model effectively integrates thinking and non-thinking modes, allowing for mode switching during conversations. Its reasoning abilities significantly surpass those of QwQ, and its general capabilities are markedly superior to Qwen2.5-Plus, reaching state-of-the-art (SOTA) levels among models of the same scale. This version is a snapshot model as of April 28, 2025.","pricing":{"cache_read":0.02252,"cache_write":0.14075,"input":0.1126,"output":1.126},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.20-beta-0309-non-reasoning","model_name":"Grok 4.20 Beta 0309","developer_id":9,"desc":"Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent tool-invocation capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, enabling it to consistently deliver precise and factual responses.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.20-beta-0309-reasoning","model_name":"Grok 4.20 Beta 0309 (reasoning)","developer_id":9,"desc":"Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent tool-invocation capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, enabling it to consistently deliver precise and factual responses.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-4.20-multi-agent-beta-0309","model_name":"Grok 4.20 Multi Agent Beta 0309","developer_id":9,"desc":"Grok 4.20 Beta is our latest flagship model, offering industry-leading speed and agent tool-invocation capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, enabling it to consistently deliver precise and factual responses.","pricing":{"cache_read":0.2,"input":2,"output":6},"types":"llm","features":"thinking,tools,function_calling,structured_outputs,long_context","input_modalities":"text,image","endpoints":"","max_output":2000000,"context_length":2000000,"schema_checked":false,"playground_checked":false},{"model_id":"o1-2024-12-17","model_name":"O1 2024 12-17","developer_id":12,"desc":"","pricing":{"cache_read":7.5,"input":15,"output":60},"types":"llm","features":"thinking","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"bai-qwen3-vl-235b-a22b-instruct","model_name":"Bai Qwen3 VL 235B A22B Instruct","developer_id":13,"desc":"The Qwen3 series open-source models include hybrid models, thinking models, and non-thinking models, with both reasoning capabilities and general abilities reaching industry SOTA levels at the same scale.","pricing":{"input":0.274,"output":1.096},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-MiniMax-M2","model_name":"CC MiniMax M2","developer_id":18,"desc":"For Claude Code only","pricing":{"input":0.1,"output":0.1},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-deepseek-v3.1","model_name":"CC DeepSeek V3.1","developer_id":7,"desc":"For Claude code only","pricing":{"input":0.56,"output":1.68},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-ernie-4.5-300b-a47b","model_name":"CC ERNIE 4.5 300B A47B","developer_id":25,"desc":"For Claude code only","pricing":{"cache_read":0,"input":0.32,"output":1.28},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-kimi-k2-instruct","model_name":"CC Kimi K2 Instruct","developer_id":15,"desc":"For Claude code only","pricing":{"input":1.1,"output":3.3},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"cc-kimi-k2-instruct-0905","model_name":"CC Kimi K2 Instruct 0905","developer_id":15,"desc":"For Claude code only","pricing":{"input":1.1,"output":3.3},"types":"llm","features":"tools,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"ahm-Phi-3-5-vision-instruct","model_name":"Ahm Phi 3.5 Vision Instruct","developer_id":3,"desc":"","pricing":{"input":0.4,"output":1.6},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"aihubmix-Cohere-command-r","model_name":"Aihubmix Cohere Command R","developer_id":6,"desc":"","pricing":{"input":0.64,"output":1.92},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"aihubmix-command-r-08-2024","model_name":"Aihubmix Command R 08 2024","developer_id":6,"desc":"","pricing":{"input":0.2,"output":0.8},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"aihubmix-command-r-plus","model_name":"Aihubmix Command R Plus","developer_id":6,"desc":"","pricing":{"input":3.84,"output":19.2},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"aihubmix-command-r-plus-08-2024","model_name":"Aihubmix Command R Plus 08 2024","developer_id":6,"desc":"","pricing":{"input":2.8,"output":11.2},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"anthropic-opus-4-6","model_name":"Anthropic Opus 4.6","developer_id":2,"desc":"Claude Opus 4.6 is Anthropic’s latest state-of-the-art reasoning model. It features an adaptive “thinking” mode that dynamically decides when to think and how much to think. At the default effort level (high), Claude will almost always engage in thinking. At lower effort levels, it may skip thinking for simple problems.\n ⚠️ The minimum cache token for claude-opus-4-6 has been increased from 1,024 to 4,096 tokens.","pricing":{"cache_read":0.5,"cache_write":6.25,"input":5,"output":25},"types":"llm","features":"thinking,tools,function_calling,structured_outputs","input_modalities":"text,image","endpoints":"","max_output":32000,"context_length":200000,"schema_checked":false,"playground_checked":false},{"model_id":"claude-3-haiku-20240229","model_name":"Claude 3 Haiku 20240229","developer_id":2,"desc":"","pricing":{"input":0.275,"output":0.275},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"claude-3-haiku-20240307","model_name":"Claude 3 Haiku 20240307","developer_id":2,"desc":"","pricing":{"input":0.275,"output":1.375},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"claude-3-haiku@20240307","model_name":"Claude 3 Haiku@20240307","developer_id":2,"desc":"","pricing":{"input":0.275,"output":1.375},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"claude-3-sonnet-20240229","model_name":"Claude 3 Sonnet 20240229","developer_id":2,"desc":"","pricing":{"input":3.3,"output":16.5},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command","model_name":"Command","developer_id":6,"desc":"","pricing":{"input":1,"output":2},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command-r","model_name":"Command R","developer_id":6,"desc":"","pricing":{"input":0.64,"output":1.92},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command-r-08-2024","model_name":"Command R 08 2024","developer_id":6,"desc":"","pricing":{"input":0.2,"output":0.8},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command-r-plus","model_name":"Command R Plus","developer_id":6,"desc":"","pricing":{"input":3.84,"output":19.2},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"command-r-plus-08-2024","model_name":"Command R Plus 08 2024","developer_id":6,"desc":"","pricing":{"input":2.8,"output":11.2},"types":"llm","features":"","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gemini-2.5-pro-exp-03-25","model_name":"Gemini 2.5 Pro","developer_id":8,"desc":"Google’s latest experimental model, highly unstable, for experience only.\nIt boasts strong reasoning and coding capabilities, able to \"think\" before responding, enhancing performance and accuracy in complex tasks. It supports multimodal inputs (text, audio, images, video) and a 1 million token context window, suitable for advanced programming, math, and science tasks.\n\nThis means Gemini 2.5 can handle more complex problems in coding, science and math, and support more context-aware agents.","pricing":{"cache_read":0.125,"input":1.25,"output":5},"types":"llm","features":"structured_outputs,tools,long_context","input_modalities":"text,image,audio,video","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-4o-mini-2024-07-18","model_name":"GPT 4o Mini 2024 07-18","developer_id":12,"desc":"","pricing":{"cache_read":0.075,"input":0.15,"output":0.6},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"gpt-oss-20b","model_name":"gpt-oss-20b","developer_id":12,"desc":"gpt-oss-20b is a 21-billion parameter open-weight model released by OpenAI under the Apache 2.0 license. Its core feature is a Mixture-of-Experts (MoE) architecture that uses only 3.6B active parameters, enabling low-latency inference and deployment on consumer GPUs. The model also supports fine-tuning, function calling, tool use, and structured outputs.","pricing":{"input":0.11,"output":0.55},"types":"llm","features":"thinking,function_calling,structured_outputs","input_modalities":"text","endpoints":"","max_output":128000,"context_length":128000,"schema_checked":false,"playground_checked":false},{"model_id":"grok-2-vision-1212","model_name":"Grok 2 Vision 1212","developer_id":9,"desc":"grok-2-vision-1212 is the latest vision model in the Grok family, delivering outstanding performance on vision-based tasks and achieving state-of-the-art results in visual mathematical reasoning and document-based question answering. It supports a wide range of visual inputs, including documents, charts, screenshots, and real-world images, making it well-suited for advanced visual understanding and reasoning use cases.\n\nThe price of calling this model in AIhubMix is ​​10% lower than on the official website.","pricing":{"input":1.8,"output":9},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"grok-vision-beta","model_name":"Grok Vision Beta","developer_id":9,"desc":"","pricing":{"input":5.6,"output":16.8},"types":"llm","features":"","input_modalities":"text,image","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"llama2-70b-4096","model_name":"Llama2 70B 4096","developer_id":11,"desc":"","pricing":{"input":0.5,"output":0.5},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"llama2-70b-40960","model_name":"Llama2 70B 40960","developer_id":11,"desc":"","pricing":{"input":0.5,"output":0.5},"types":"llm","features":"","input_modalities":"","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-turbo","model_name":"Qwen Turbo","developer_id":13,"desc":"","pricing":{"cache_read":0.0092,"input":0.046,"output":0.092},"types":"llm","features":"long_context","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false},{"model_id":"qwen-turbo-2024-11-01","model_name":"Qwen Turbo 2024 11-01","developer_id":13,"desc":"","pricing":{"input":0.046,"output":0.092},"types":"llm","features":"long_context","input_modalities":"text","endpoints":"","max_output":0,"context_length":0,"schema_checked":false,"playground_checked":false}],"message":"","success":true}