How to Calculate LLM GPU VRAM Requirements for Production Inference
Running open-source large language models (such as DeepSeek-R1, Meta Llama 3.3 70B, Qwen 2.5, or Mistral) locally or in cloud production requires accurately budgeting GPU memory. Running out of VRAM leads directly to CUDA out of memory crashes or severe CPU offloading bottlenecks.
Total inference VRAM is governed by three independent components:
Total VRAM (GB) = Model Weights Memory + KV Cache Memory + CUDA Runtime Overhead1. Model Weights Memory Formula
The base memory required just to hold the neural network parameters in VRAM depends solely on the parameter count and the quantization bit-depth:
- 16-bit (FP16 / BF16):
Parameters (Billions) × 2.0 Bytes. Example: Llama-3-70B requires ~140 GB. - 8-bit (FP8 / Q8_0 GGUF):
Parameters (Billions) × 1.0 Bytes. Example: 70B requires ~70 GB. - 4-bit (AWQ / GPTQ / Q4_K_M GGUF):
Parameters (Billions) × 0.55 Bytes. Example: 70B requires ~38.5 GB. - 2-bit (Q2_K / EXL2):
Parameters (Billions) × 0.35 Bytes. Example: 70B requires ~24.5 GB.
2. KV Cache Scaling & Attention Mechanisms
As the conversation or document length grows, the attention Key-Value (KV) cache grows linearly with context tokens and batch size:
KV Cache Bytes = 2 × num_layers × num_kv_heads × head_dim × bytes_per_element × context_length × batch_sizeModern models utilize Grouped-Query Attention (GQA) (such as Llama 3 with 8 KV heads vs 64 Query heads), reducing the KV cache memory footprint by 8x compared to legacy Multi-Head Attention (MHA). Furthermore, DeepSeek-V3 and R1 utilize Multi-Head Latent Attention (MLA), which compresses key-value vectors into a low-dimensional latent space (576 dimensions), allowing extreme 128k context scaling with minimal VRAM overhead.
3. Reference VRAM Table for Popular Open-Source LLMs
| Model | Parameters | 4-bit (Q4_K_M) | 8-bit (FP8) | 16-bit (FP16) | Recommended GPU |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8.03B | 5.8 GB | 9.6 GB | 17.5 GB | 1x RTX 3060 12GB |
| Qwen 2.5 14B | 14.7B | 9.8 GB | 16.5 GB | 31.2 GB | 1x RTX 4080 16GB |
| Llama 3.3 70B | 70.6B | 41.5 GB | 76.0 GB | 148.0 GB | 2x RTX 3090 (48GB) |
| DeepSeek-R1 | 671B MoE | 380.0 GB | 710.0 GB | 1,380 GB | 8x H100 80GB (640GB) |