Gemini 2.5 Flash Lite
Google logo

Gemini 2.5 Flash Lite

gemini-2.5-flash-litellms.txt
Google
Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

Pricing

PricingCache ReadInput AudioInput Audio Cached Web SearchCache Storage
$0.100$0.400
$0.01/M tokens$0.3/M tokens$0.03/M tokens$0.035/request$1/h/M tokens

Input Modalities

  • Text
  • Vision
  • Audio
  • Video

Output Modalities

  • Text

Capabilities

  • Tools
  • Tool calling
  • Structured outputs
  • Long context

Providers

VertexAI gemini-2.5-flash-lite
Pricing$0.100$0.400
Cache Read$0.01/M tokens
Input Audio$0.3/M tokens
Input Audio Cached $0.03/M tokens
Web Search$0.035/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency2.9S
Throughput40.5TPS
Uptime
99.41% uptime 2 days ago
98.69% uptime yesterday
99.38% uptime today
Google AI Studio gemini-2.5-flash-lite
Pricing$0.100$0.400
Cache Read$0.01/M tokens
Input Audio$0.3/M tokens
Input Audio Cached $0.03/M tokens
Web Search$0.035/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency1.2S
Throughput24.6TPS
Uptime
98.96% uptime 2 days ago
98.88% uptime yesterday
96.78% uptime today

Performance for gemini-2.5-flash-lite

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="gemini-2.5-flash-lite",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Gemini 2.5 Flash Lite?

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.