by Nvidia
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid Mamba-Transformer architecture. Activating just 12B parameters, it delivers maximum compute efficiency and accuracy for complex multi-agent applications. Additionally, it features an expansive context length of 262,144 tokens to handle large-scale inputs seamlessly.
| Context window | 262,144 tokens |
| Modalities | text |
| Features | reasoning, tool_calling, long_context |
| Endpoints | chat_completions |
Use nemotron-3-super-120b-a12b-free via the AIHubMix unified API — one interface for every major LLM.