Text

Nemotron Nano 9B v2

NVIDIA's efficient 9B hybrid architecture model with Mamba-2 and attention layers

Command to Run on Jetson Benchmark Model Details

Parameters 9B

Modalities

Text

Context Length 128K

License NVIDIA Open Model License

Precision

NVFP4

Serve the model

Start server

Choose module, then engine and optional parameters on the left, then copy the serve command by clicking the button on the right.

Command

Call the model over Web API

Copy a client command below and paste it into your terminal to make a Web API request to the model you just served.

curl -s http://${JETSON_HOST}:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/NVIDIA-Nemotron-Nano-9B-v2-NVFP4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

from openai import OpenAI

client = OpenAI(
    base_url="http://${JETSON_HOST}:8000/v1",
    api_key="not-needed",  # vLLM / llama.cpp typically do not enforce a key
)

completion = client.chat.completions.create(
    model="nvidia/NVIDIA-Nemotron-Nano-9B-v2-NVFP4",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(completion.choices[0].message.content)

Benchmark

Nemotron Nano 9B V2 · vLLM · NVFP4 / W4A16 · ISL 2048 / OSL 128

Engine

Concurrency

C = concurrent requests. Results will vary with image, clocks, and workload.

Model Details

View on HuggingFace

NVIDIA Nemotron Nano 9B v2 is a quantized large language model trained from scratch by NVIDIA, designed as a unified model for both reasoning and non-reasoning tasks. It generates a reasoning trace before concluding with a final response, with configurable reasoning via system prompt.

Architecture

The model uses a hybrid architecture:

56 layers total: 27 Mamba layers, 25 MLP layers, 4 attention layers
NVFP4 quantization with Mamba and MLP layers quantized
Attention layers and Conv1d components kept in BF16 for accuracy
Quantization-Aware Distillation (QAD) applied for accuracy recovery

Inputs and Outputs

Input: Text

Output: Text

Intended Use Cases

AI Agent Systems: Autonomous agents with reasoning capabilities
Chatbots: General purpose conversational AI
RAG Systems: Retrieval-augmented generation applications
Instruction Following: General instruction-following tasks
Code Generation: Programming assistance in multiple languages

Supported Languages

English, German, Spanish, French, Italian, Japanese, and coding languages.

This model is ready for commercial use.