meta-llama/Llama-3.1-8B-Instruct
meta-llama/Llama-3.1-8B-Instruct
Meta Llama 3.1 8B, the default open model baseline.
Parameters
8B
Context
128K
Downloads
6.1M
Likes
8K
Architecture
β
GPU Requirements
VRAM / RAM / disk estimates by quantization and usage scenario
| Quantization | Minimum | Recommended | Production |
|---|---|---|---|
| FP1616bit | 23.8 GB RAM 35.7 Β· Disk 62.3 | 40.4 GB RAM 60.6 Β· Disk 88.3 | 138.2 GB RAM 207.3 Β· Disk 127.3 |
| FP88bit | 14.4 GB RAM 32.0 Β· Disk 50.6 | 30.6 GB RAM 45.8 Β· Disk 76.6 | 127.9 GB RAM 191.8 Β· Disk 115.6 |
| INT88bit | 14.4 GB RAM 32.0 Β· Disk 50.6 | 30.6 GB RAM 45.8 Β· Disk 76.6 | 127.9 GB RAM 191.8 Β· Disk 115.6 |
| AWQ4bit | 9.8 GB RAM 32.0 Β· Disk 45.0 | 25.8 GB RAM 38.7 Β· Disk 71.0 | 122.9 GB RAM 184.4 Β· Disk 110.0 |
| GPTQ4bit | 9.9 GB RAM 32.0 Β· Disk 45.1 | 25.9 GB RAM 38.8 Β· Disk 71.1 | 123.0 GB RAM 184.5 Β· Disk 110.1 |
| GGUF4bit | 10.0 GB RAM 32.0 Β· Disk 45.3 | 26.0 GB RAM 39.0 Β· Disk 71.3 | 123.1 GB RAM 184.7 Β· Disk 110.3 |
GPU Recommendations
1Γ RTX 5090
32 GB VRAM
β 464.2 tok/s
Effective production throughput β 324.9 tok/s
score 0.6818
2Γ Tesla V100
32 GB VRAM
β 0 tok/s (N/A β benchmark unavailable)
Effective production throughput β 0 tok/s
score 0.6315
8Γ H200 SXM
1128 GB VRAM
β 3979.3 tok/s
Effective production throughput β 2785.5 tok/s
score 0.2197
1Γ RTX A6000
48 GB VRAM
β 199 tok/s
Effective production throughput β 139.3 tok/s
score 0.6334
π¬ Community β real deployment experience
0 Articles Β· 0 Benchmarks
GPUs that fit (single card)
RTX A6000
48 GB VRAM
$264
/mo Β· vast
RTX 5090
32 GB VRAM
$282
/mo Β· vast
RTX 5090D
32 GB VRAM
$294
/mo Β· vast
L40S
48 GB VRAM
$341
/mo Β· vast
A100 PCIE
80 GB VRAM
$400
/mo Β· vast