NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring an extensive 1,000,000-token context length. Built on a hybrid Transformer-Mamba mixture-of-experts (MoE) architecture, it operates with 55B active parameters out of a total of 550B parameters. This advanced structure is designed to support sophisticated reasoning and complex task orchestration.
Pricing
- Input Tokens: $0.000 /M tokens
- Output Tokens: $0.000 /M tokens
- Cache Read: $0.000 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Capabilities
- Long context
Try this model
Python
