Llama Models
Usage
5.1B
374K
22 of 52
Which models that traffic went to
- llama-3.3-70b34.2%1.8B
- llama-4-maverick29.7%1.5B
- llama-4-scout24.2%1.2B
- llama3.1-8b5.7%294M
- llama-3.1-70b4.7%241M
- deepinfra-llama-3.1-8b-instant1.1%57.7M
- deepinfra-llama-3.3-70b-instant-turbo0.4%19.5M
- groq-llama-3.3-70b-versatile<0.1%1.5M
- 5 more models<0.1%137K
- llama-3.3-70b41.8%156K
- llama-4-scout31.9%119K
- deepinfra-llama-3.1-8b-instant12.0%44.9K
- llama-4-maverick7.9%29.6K
- llama3.1-8b4.3%16.1K
- llama-3.1-70b1.4%5.1K
- deepinfra-llama-3.3-70b-instant-turbo0.7%2.6K
- groq-llama-3.3-70b-versatile<0.1%77
- 14 more models<0.1%76
All 52 Llama Models
Open in model listLlama on AIHubMix
Which Llama model should I start with?
llama-4-maverick at $0.20/M input — the cheapest entry here that declares tool calling, and it carries a 1.05M context. Move up to aihubmix-Llama-3-1-405B-Instruct when answer quality matters more than cost.
Why are there several entries for the same model?
Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (meta-llama/…), and some differ only in capitalisation, kept so older integrations keep working.
The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.
Do I need a separate Llama account?
No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.
Start calling Llama in one line
One key, one endpoint, 844 models across 35 providers.
