NVIDIA Nemotron 3 Super

nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16

Parameters

120B

Context

1024K

Downloads

818.9K

Likes

415

Architecture

moe_transformer

View on HuggingFace

GPU Requirements

VRAM / RAM / disk estimates by quantization and usage scenario

QuantizationMinimumRecommendedProduction
FP1616bit
287.7 GB
RAM 431.6 · Disk 387.7
316.3 GB
RAM 474.5 · Disk 413.7
426.1 GB
RAM 639.1 · Disk 452.7
FP88bit
146.3 GB
RAM 219.5 · Disk 213.3
168.5 GB
RAM 252.8 · Disk 239.3
271.8 GB
RAM 407.7 · Disk 278.3
INT88bit
146.3 GB
RAM 219.5 · Disk 213.3
168.5 GB
RAM 252.8 · Disk 239.3
271.8 GB
RAM 407.7 · Disk 278.3
AWQ4bit
78.3 GB
RAM 117.4 · Disk 129.4
97.4 GB
RAM 146.1 · Disk 155.4
197.6 GB
RAM 296.4 · Disk 194.4
GPTQ4bit
79.2 GB
RAM 118.8 · Disk 130.5
98.3 GB
RAM 147.4 · Disk 156.5
198.6 GB
RAM 297.9 · Disk 195.5
GGUF4bit
80.9 GB
RAM 121.4 · Disk 132.7
100.1 GB
RAM 150.2 · Disk 158.7
200.5 GB
RAM 300.8 · Disk 197.7

GPU Recommendations

Best ValueBEST

4× H200 SXM

564 GB VRAM

$10,220/mo

≈ 198.7 tok/s

Effective production throughput ≈ 139.1 tok/s

score 0.3154

Cheapest

8× Tesla V100

128 GB VRAM

$162/mo

≈ 0 tok/s (N/A — benchmark unavailable)

Effective production throughput ≈ 0 tok/s

score 0.5538

Deploy on vast ↗
Performance

8× H200 SXM

1128 GB VRAM

$20,440/mo

≈ 265 tok/s

Effective production throughput ≈ 185.5 tok/s

score 0.2197

Min Complexity

1× H200 SXM

141 GB VRAM

$2,555/mo

≈ 82.8 tok/s

Effective production throughput ≈ 58 tok/s

score 0.6552

💬 Community — real deployment experience

0 Articles · 0 Benchmarks

View Community →

GPUs that fit (single card)

H200 SXM

141 GB VRAM

$2,555

/mo · lambda

B200

180 GB VRAM

$3,013

/mo · vast

B200 SXM

180 GB VRAM

AmciHub

AI Model Deployment & Compute Intelligence — analyze AI model requirements, GPU performance, cloud pricing and deployment costs to find the right deployment solution.

© 2026 AmciHub. All rights reserved. Model → Requirement → GPU → Benchmark → Cloud → Cost → Recommendation