Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost. Designed for the agentic era, it excels at code generation, agentic execution, and spatial reasoning. This model is particularly effective for rapid agentic loops involving complex coding cycles and iterations.
Pricing
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Context length
- 1.05M tokens
Max output
- 65.5K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Providers
VertexAI gemini-3.6-flash
Pricing$0.750$3.750
Cache Read$0.075/M tokens
Input Video$0.75/M tokens
Input Audio$0.75/M tokens
Web Search$0.014/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency7.1S
Throughput96.9TPS
Uptime
98.41% uptime 2 days ago
99.58% uptime yesterday
99.90% uptime today
Google AI Studio gemini-3.6-flash
Pricing$0.750$3.750
Cache Read$0.075/M tokens
Input Video$0.75/M tokens
Input Audio$0.75/M tokens
Web Search$0.014/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency5.4S
Throughput84.0TPS
Uptime
94.08% uptime 2 days ago
97.68% uptime yesterday
97.42% uptime today
Performance for gemini-3.6-flash
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
