Gemma 3 4B It
Google logo

Gemma 3 4B It

gemma-3-4b-it
Google
Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned variants. Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions. Gemma 3 models are well-suited for a variety of text generation and image understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone.

Pricing

  • Input Tokens: $0.200 /M tokens
  • Output Tokens: $0.200 /M tokens
  • Cache Read: $0.000 /M tokens

Input Modalities

    Providers

    Deepinfra deepinfra-gemma-3-4b-it
    Pricing$0.044$0.088
    Context0
    Max output0
    Latency1.3S
    Throughput20.2TPS
    Uptime
    100.00% uptime 2 days ago
    100.00% uptime yesterday
    100.00% uptime today
    Google AI Studio google-gemma-3-4b-it
    Pricing$0.200$0.200
    Cache$0.000
    Context0
    Max output0
    Latency-
    Throughput-
    Uptime
    0.00% uptime 2 days ago
    0.00% uptime yesterday
    0.00% uptime today

    Performance for gemma-3-4b-it

    Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

    Uptime
    Loading...
    Latency
    Loading...
    Throughput
    Loading...

    Try this model

    Python
    import os
    from openai import OpenAI
    
    client = OpenAI(
        api_key=os.environ["AIHUBMIX_API_KEY"],
        base_url="https://aihubmix.com/v1",
    )
    
    response = client.chat.completions.create(
        model="gemma-3-4b-it",
        messages=[
          {
            "role": "user",
            "content": "Hello, how are you?"
          }
        ],
        max_tokens=1024,
        stream=False,
    )
    
    print(response.choices[0].message.content)

    Frequently asked questions

    What is Gemma 3 4B It?

    Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned variants. Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions. Gemma 3 models are well-suited for a variety of text generation and image understanding tasks, including question answering, summarization, and reasoning. Their relatively small size makes it possible to deploy them in environments with limited resources such as laptops, desktops or your own cloud infrastructure, democratizing access to state of the art AI models and helping foster innovation for everyone.