Gemini 2.5 Pro Preview 06-05
Google logo

Gemini 2.5 Pro Preview 06-05

gemini-2.5-pro-preview-06-05
Google
Google’s latest multimodal flagship model, combining exceptional coding and reasoning capabilities. Its massive 1 million token context window (soon to expand to 2 million) places it at the top of the WebDevArena and LMArena leaderboards. It is particularly well-suited for developing aesthetically pleasing and highly functional interactive web applications, code transformation, and complex workflows. The newly introduced "reasoning budget" feature cleverly balances cost and performance, while optimized tool calls and response styles further enhance development efficiency, making it the ideal choice for rapid prototyping and advanced coding.

Pricing

TierPricingCache ReadWeb SearchCache Storage
Input<=200K
$1.250$10.000
$0.125/M tokens$0.035/request$4.5/h/M tokens
200K<Input
$2.500$15.000
$0.25/M tokens$0.035/request$4.5/h/M tokens

Input Modalities

  • Text
  • Vision
  • Audio
  • Video

Output Modalities

  • Text

Capabilities

  • Thinking
  • Tools
  • Tool calling
  • Structured outputs
  • Long context

Providers

Google AI Studio gemini-2.5-pro-preview-06-05
Pricing$1.250$10.000
Cache Read$0.125/M tokens
Web Search$0.035/request
Cache Storage$4.5/h/M tokens
Pricing$2.500$15.000
Cache Read$0.25/M tokens
Web Search$0.035/request
Cache Storage$4.5/h/M tokens
Context1M
Max output65K
Latency8.0S
Throughput78.2TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
VertexAI gemini-2.5-pro-preview-06-05
Pricing$1.250$10.000
Cache Read$0.125/M tokens
Web Search$0.035/request
Cache Storage$4.5/h/M tokens
Pricing$2.500$15.000
Cache Read$0.25/M tokens
Web Search$0.035/request
Cache Storage$4.5/h/M tokens
Context1M
Max output65K
Latency8.0S
Throughput78.2TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today

Performance for gemini-2.5-pro-preview-06-05

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="gemini-2.5-pro-preview-06-05",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Gemini 2.5 Pro Preview 06-05?

Google’s latest multimodal flagship model, combining exceptional coding and reasoning capabilities. Its massive 1 million token context window (soon to expand to 2 million) places it at the top of the WebDevArena and LMArena leaderboards. It is particularly well-suited for developing aesthetically pleasing and highly functional interactive web applications, code transformation, and complex workflows. The newly introduced "reasoning budget" feature cleverly balances cost and performance, while optimized tool calls and response styles further enhance development efficiency, making it the ideal choice for rapid prototyping and advanced coding.