by Nvidia
Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It features a powerful 128,000-token context length to handle extensive visual and textual inputs. The model utilizes a hybrid Transformer-Mamba architecture, successfully combining transformer-level accuracy with Mamba's structural advantages.
| Context window | 128,000 tokens |
| Modalities | text, image |
| Features | reasoning, tool_calling, long_context |
| Endpoints | chat_completions |
Use nemotron-nano-12b-v2-vl-free via the AIHubMix unified API — one interface for every major LLM.