Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring 400B total parameters and 128 experts, while activating an efficient 17B parameters per inference. Engineered for peak performance, it excels at advanced multimodal tasks.
Maverick natively supports text and image input, producing multilingual text and code. With a 1-million-token context window and instruction tuning, it is optimized for complex image reasoning and general-purpose assistant-like interactions.
Released under the Llama 4 Community License, Maverick is ideal for research and commercial applications demanding state-of-the-art multimodal understanding and high throughput.
Pricing
- Input Tokens: $0.200 /M tokens
- Output Tokens: $0.200 /M tokens
Input Modalities
- Text
- Vision
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
Providers
Baidu bai-llama-4-maverick-17b-128e-instruct
Pricing$0.960$2.880
Context1M
Max output32K
Latency1.6S
Throughput77.3TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Groq groq-llama-4-maverick-17b-128e-instruct
Pricing$0.220$0.660
Context131K
Max output8K
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Deepinfra deepinfra-llama-4-maverick-17b-128e-instruct
Pricing$0.330$1.320
Context1M
Max output128K
Latency0.4S
Throughput45.8TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Chutes chutesai/Llama-4-Maverick-17B-128E-Instruct-FP8
Pricing$0.250$0.250
Cache$0.000
Context1M
Max output16K
Latency0.2S
Throughput97.8TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Performance for llama-4-maverick
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
