by Nvidia
NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA. Designed as a unified model, it efficiently handles both reasoning and non-reasoning tasks to respond to a wide range of user queries. With a generous context length of 128,000 tokens, it is highly capable of processing long and complex documents.
| Context window | 128,000 tokens |
| Modalities | text |
| Features | reasoning, tool_calling, long_context |
| Endpoints | chat_completions |
Use nemotron-nano-9b-v2-free via the AIHubMix unified API — one interface for every major LLM.