NVIDIA Nemotron 3 Ultra

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

Parameters

550B

Context

1024K

Downloads

497K

Likes

316

Architecture

moe_transformer

View on HuggingFace

GPU Requirements

VRAM / RAM / disk estimates by quantization and usage scenario

QuantizationMinimumRecommendedProduction
FP1616bit
1300.9 GB
RAM 1951.3 · Disk 1637.2
1375.5 GB
RAM 2063.3 · Disk 1663.2
1531.4 GB
RAM 2297.0 · Disk 1702.2
FP88bit
652.9 GB
RAM 979.4 · Disk 838.1
698.1 GB
RAM 1047.2 · Disk 864.1
824.5 GB
RAM 1236.7 · Disk 903.1
INT88bit
652.9 GB
RAM 979.4 · Disk 838.1
698.1 GB
RAM 1047.2 · Disk 864.1
824.5 GB
RAM 1236.7 · Disk 903.1
AWQ4bit
341.1 GB
RAM 511.6 · Disk 453.5
372.1 GB
RAM 558.2 · Disk 479.5
484.3 GB
RAM 726.4 · Disk 518.5
GPTQ4bit
345.1 GB
RAM 517.7 · Disk 458.5
376.4 GB
RAM 564.5 · Disk 484.5
488.7 GB
RAM 733.1 · Disk 523.5
GGUF4bit
353.2 GB
RAM 529.9 · Disk 468.5
384.8 GB
RAM 577.2 · Disk 494.5
497.5 GB
RAM 746.3 · Disk 533.5

GPU Recommendations

Best ValueBEST

8× H200 SXM

1128 GB VRAM

$20,440/mo

≈ 57.8 tok/s

Effective production throughput ≈ 40.5 tok/s

score 0.3077

Cheapest

8× RTX A6000

384 GB VRAM

$1,678/mo

≈ 9.2 tok/s

Effective production throughput ≈ 6.4 tok/s

score 0.4893

Deploy on vast ↗
Performance

4× H200 SXM

564 GB VRAM

$10,220/mo

≈ 43.4 tok/s

Effective production throughput ≈ 30.4 tok/s

score 0.4453

Min Complexity

4× H200 SXM

564 GB VRAM

$10,220/mo

≈ 43.4 tok/s

Effective production throughput ≈ 30.4 tok/s

score 0.4453

💬 Community — real deployment experience

0 Articles · 0 Benchmarks

View Community →
AmciHub

AI Model Deployment & Compute Intelligence — analyze AI model requirements, GPU performance, cloud pricing and deployment costs to find the right deployment solution.

© 2026 AmciHub. All rights reserved. Model → Requirement → GPU → Benchmark → Cloud → Cost → Recommendation