Pricing
| Pricing | Cache Read | Input Video | Input Audio | Web Search | Cache Storage |
|---|---|---|---|---|---|
Pricing | Cache Read | Input Video | Input Audio | Web Search | Cache Storage |
$0.750$3.750 | $0.075/M tokens | $0.75/M tokens | $0.75/M tokens | $0.014/request | $1/h/M tokens |
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Capabilities
- Thinking
- Web
- Tools
- Tool calling
- Structured outputs
- Long context
Try this model
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
response = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[
{
"role": "user",
"content": "Hello, how are you?"
}
],
max_tokens=1024,
stream=False,
)
print(response.choices[0].message.content)Frequently asked questions
What is gemini-3.7-flash?
What is the context length of gemini-3.7-flash?
How much does gemini-3.7-flash cost?
What modalities does gemini-3.7-flash support?
What capabilities does gemini-3.7-flash support?
How do I call gemini-3.7-flash via API?
Who created gemini-3.7-flash?
Compare gemini-3.7-flash
More models from Google
- Input: $ 0 /M
- Output: $ 0 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.7 Flash free version: fFree model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version: gemini-3.7-flash
- Input: $ 1.5 /M
- Output: $ 7.5 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $1/M tokens
- Input Video: $1/M tokens
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at code generation, agentic execution, and spatial reasoning. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.
Input price: $0.25/M, Text output: $1.5/M, Image output: $30/M
Google's newest, most compact, and most cost-effective image generation and editing model, designed for large-scale use.
- Input: $ 0.3 /M
- Output: $ 2.499999 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $1/M tokens
- Input Video: $1/M tokens
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.
- Input: $ 0 /M
- Output: $ 0 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version:gemini-3.5-flash-lite.
- Input: $ 0 /M
- Output: $ 0 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.6 Flash free version: fFree model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version: gemini-3.6-flash
AIHubMix© 2023 - 2026 AIHubMix, LLC