Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned variants. Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions. Gemma 3 models are well-suited for a variety of text generation and image understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone.
Pricing
- Input Tokens: $0.200 /M tokens
- Output Tokens: $0.200 /M tokens
- Cache Read: $0.000 /M tokens
Input Modalities
Providers
Deepinfra deepinfra-gemma-3-4b-it
Pricing$0.044$0.088
Context0
Max output0
Latency1.3S
Throughput20.2TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Google AI Studio google-gemma-3-4b-it
Pricing$0.200$0.200
Cache$0.000
Context0
Max output0
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Performance for gemma-3-4b-it
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
