Pricing
| Pricing | Input Video | Input Audio | Web Search | Cache Storage |
|---|---|---|---|---|
Pricing | Input Video | Input Audio | Web Search | Cache Storage |
$0.250$1.500 | $0.25/M tokens | $0.5/M tokens | $0.014/request | $1/h/M tokens |
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Context length
- 1.05M tokens
Max output
- 65.5K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Try this model
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
response = client.chat.completions.create(
model="gemini-3.1-flash-lite-preview",
messages=[
{
"role": "user",
"content": "Hello, how are you?"
}
],
max_tokens=1024,
stream=False,
)
print(response.choices[0].message.content)Frequently asked questions
What is gemini-3.1-flash-lite-preview?
What is the context length of gemini-3.1-flash-lite-preview?
How much does gemini-3.1-flash-lite-preview cost?
What modalities does gemini-3.1-flash-lite-preview support?
What capabilities does gemini-3.1-flash-lite-preview support?
How do I call gemini-3.1-flash-lite-preview via API?
Who created gemini-3.1-flash-lite-preview?
More models from Google
- Input: $ 1.5 /M
- Output: $ 7.5 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $1/M tokens
- Input Video: $1/M tokens
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at code generation, agentic execution, and spatial reasoning. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.
Input price: $0.25/M, Text output: $1.5/M, Image output: $30/M
Google's newest, most compact, and most cost-effective image generation and editing model, designed for large-scale use.
- Input: $ 0.3 /M
- Output: $ 2.499999 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $1/M tokens
- Input Video: $1/M tokens
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.
- Input: $ 0 /M
- Output: $ 0 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version:gemini-3.5-flash-lite.
- Input: $ 0 /M
- Output: $ 0 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.6 Flash free version: fFree model resources are limited and provided only for trial use; stability cannot be guaranteed, and you may encounter 429 errors during use. If you need to use it in a production environment and require unlimited concurrency with absolute stability, please choose the official version: gemini-3.6-flash
- Input: $ 1.5 /M
- Output: $ 9 /M
- Web Search: $0.014/request
- Cache Storage: $1/h/M tokens
- Input Audio: $2/M tokens
- Input Video: $1/M tokens
Gemini 3.5 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at sub-agent deployment, multi-step workflows, and long-horizon tasks at scale. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.
AIHubMix© 2023 - 2026 AIHubMix, LLC