Calculate GPU count and cost for LLM training & inference
This calculator uses NVIDIA official benchmarks (NIM, ๐ฅ MLPerfโข v6.1, TensorRT-LLM) and ๐ MLCommons Standards.
Tier 1: Megatron-Bridge official โ LLaMA3-8B: micro_batch=2, 405B: micro_batch=1. Global batch = micro ร GPUs ร grad_accum.
All model weights updated. Highest memory usage.
Larger r = more trainable params = more memory. Default r=16 is standard for most tasks.
Typical: 1โ5 epochs
Typical: 75โ90%
Longer context = more activation memory
Reduces per-GPU memory. Requires NVLink for TP4+
Splits model layers across GPU groups. For large models (100B+) on multi-node clusters.
More micro-batches = less pipeline bubble. Used by Megatron-LM / NeMo. (Narayanan et al. 2021, NVIDIA SC'21)
Uses curated NVIDIA NIM, MLPerf, and TensorRT-LLM benchmark data โ not just model-size estimates.
Calculates training time and GPU count from FLOPS, MFU, and parallel efficiency โ not just VRAM sizing.
Shows transparent source badges so you know whether results are benchmark-based or estimated.
Separates MoE weight memory from active compute for accurate sizing of DeepSeek, Mixtral, and Qwen models.
Grouped by architecture and precision support. Sorted by performance within each group.
| GPU | VRAM | FP16 TFLOPS | Mem BW | FP16 | FP8 | FP4 | $/hr (est.) | Architecture |
|---|---|---|---|---|---|---|---|---|
| โก Blackwell Ultra โ FP16 / FP8 / FP4 | ||||||||
| GB300 NVL72 ๐ Top | 288 GB/GPU 20.7 TB system |
7,200 | 10,000 GB/s | โ | โ | โ | ~$15.00 | Blackwell Ultra |
| ๐ฅ Blackwell โ FP16 / FP8 / FP4 | ||||||||
| GB200 NVL72 | 192 GB/GPU 13.4 TB system |
5,000 | 8,000 GB/s | โ | โ | โ | ~$10.50 | Blackwell |
| B300 DGX (8-GPU) ๐ MLPerf v6.1 | 288 GB/GPU 2.1 TB system |
4,500 | 8,000 GB/s | โ | โ | โ | ~$10.75* | Blackwell Ultra |
| B200 DGX (8-GPU) | 180 GB/GPU 1.4 TB system |
4,500 | 8,000 GB/s | โ | โ | โ | ~$8.00 | Blackwell |
| RTX PRO 6000 Blackwell | 96 GB | 234 | 1,597 GB/s | โ | โ | โ | ~$2.50 | Blackwell |
| Hopper โ FP16 / FP8 | ||||||||
| H200 SXM | 141 GB | 1,979 | 4,800 GB/s | โ | โ | โ | ~$4.50 | Hopper |
| H100 SXM (default) | 80 GB | 1,979 | 3,350 GB/s | โ | โ | โ | ~$2.50 | Hopper |
| H100 PCIe | 80 GB | 1,513 | 2,000 GB/s | โ | โ | โ | ~$2.00 | Hopper |
| AMD CDNA4 โ FP16 / FP8 | ||||||||
| AMD Instinct MI355X MLPerf v6.0 | 288 GB | 3,840 | 8,000 GB/s | โ | โ | โ | ~$8.00 | CDNA4 |
| Ada Lovelace โ FP16 / FP8 | ||||||||
| L40S | 48 GB | 733 | 864 GB/s | โ | โ | โ | ~$0.80 | Ada |
| L4 | 24 GB | 242 | 300 GB/s | โ | โ | โ | ~$0.35 | Ada |
| Ampere โ FP16 only | ||||||||
| A100 80GB SXM | 80 GB | 1,248 | 2,000 GB/s | โ | โ | โ | ~$1.80 | Ampere |
| A100 40GB | 40 GB | 1,248 | 1,555 GB/s | โ | โ | โ | ~$1.20 | Ampere |
* Prices are cloud on-demand estimates and vary by provider. GB300/GB200 TFLOPS and bandwidth are per-GPU figures; total system performance scales with 72 GPUs. Source: NVIDIA official specs + MLPerf v6.1.