nemotron-nano-12b-v2-vl-free

by Nvidia

Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It features a powerful 128,000-token context length to handle extensive visual and textual inputs. The model utilizes a hybrid Transformer-Mamba architecture, successfully combining transformer-level accuracy with Mamba's structural advantages.

Specifications

Context window128,000 tokens
Modalitiestext, image
Featuresreasoning, tool_calling, long_context
Endpointschat_completions

More from Nvidia

Use nemotron-nano-12b-v2-vl-free via the AIHubMix unified API — one interface for every major LLM.