LLM VRAM Calculator
Pick a local model, a quantization and a context length to see how much GPU memory it needs, and which graphics cards and Macs it fits on.
- Model file 4.8 GB listed as 5.2 GB
- Context 0.6 GB
- Allowance 0.5 GB
Graphics cards
- 8 GB VRAMGeForce RTX 5060 Ti 8 GBGeForce RTX 5060GeForce RTX 5050 +2 more 2.1 GB spare
- 12 GB VRAMGeForce RTX 5070GeForce RTX 4070 TiGeForce RTX 4070 SUPER +1 more 6.1 GB spare
- 16 GB VRAMGeForce RTX 5080GeForce RTX 5070 TiGeForce RTX 5060 Ti 16 GB +4 more 10 GB spare
- 24 GB VRAMGeForce RTX 4090 18 GB spare
- 32 GB VRAMGeForce RTX 5090 26 GB spare
Macs
- 16 GB 12 GB usableMacBook Air (M5), 16 GBMac mini (M6), 16 GB 6.1 GB spare
- 24 GB 18 GB usable
- 32 GB 24 GB usableMacBook Air (M5), 32 GBMac mini (M6), 32 GB 18 GB spare
- 48 GB 36 GB usableMac mini (M5 Pro), 48 GB 30 GB spare
- 64 GB 48 GB usableMac mini (M5 Pro), 64 GB 42 GB spare
| Model | Quantization | File | At 4K context | At max context | Smallest card |
|---|---|---|---|---|---|
| Qwen3 0.6B | Q4_K_M | 0.523 GB | 1.4 GB | 5.4 GB | 8 GB VRAM |
| Qwen3 1.7B | Q4_K_M | 1.4 GB | 2.2 GB | 6.2 GB | 8 GB VRAM |
| Qwen3 4B | Q4_K_M | 2.6 GB | 3.5 GB | 8.5 GB | 8 GB VRAM |
| Qwen3 8B | Q4_K_M | 5.2 GB | 5.9 GB | 11 GB | 8 GB VRAM |
| Q8_0 | 8.9 GB | 9.4 GB | 14 GB | 12 GB VRAM | |
| Qwen3 14B | Q4_K_M | 9.3 GB | 9.8 GB | 15 GB | 12 GB VRAM |
| Q8_0 | 16 GB | 16 GB | 22 GB | 24 GB VRAM | |
| Qwen3 32B | Q4_K_M | 20 GB | 20 GB | 29 GB | 24 GB VRAM |
| Q8_0 | 35 GB | 34 GB | 43 GB | Larger than we track | |
| Qwen2.5 0.5B | Q4_K_M | 0.398 GB | 0.9 GB | 1.2 GB | 8 GB VRAM |
| Qwen2.5 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Qwen2.5 72B | Q4_K_M | 47 GB | 46 GB | 54 GB | Larger than we track |
| Qwen2.5 Coder 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 Coder 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 Coder 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 Coder 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 Coder 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Phi-4 14B | Q4_K_M | 9.1 GB | 9.8 GB | 12 GB | 12 GB VRAM |
| Phi-4 mini 3.8B | Q4_K_M | 2.5 GB | 3.3 GB | 19 GB | 8 GB VRAM |
| Phi-3 mini 3.8B | Q4_K_M | 2.4 GB | 4.2 GB | 4.2 GB | 8 GB VRAM |
| Phi-3 medium 14B | Q4_K_M | 8.6 GB | 9.3 GB | 34 GB | 12 GB VRAM |
| Mistral 7B | Q4_K_M | 4.4 GB | 5.1 GB | 8.6 GB | 8 GB VRAM |
| Mistral Nemo 12B | Q4_K_M | 7.5 GB | 8.1 GB | 27 GB | 12 GB VRAM |
| Mistral Small 22B | Q4_K_M | 13 GB | 13 GB | 20 GB | 16 GB VRAM |
| Mistral Small 24B | Q4_K_M | 14 GB | 14 GB | 19 GB | 16 GB VRAM |
| DeepSeek-R1 Distill 1.5B | Q4_K_M | 1.1 GB | 1.6 GB | 5 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 7B | Q4_K_M | 4.7 GB | 5.1 GB | 12 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 14B | Q4_K_M | 9 GB | 9.6 GB | 33 GB | 12 GB VRAM |
| DeepSeek-R1 Distill 32B | Q4_K_M | 20 GB | 20 GB | 51 GB | 24 GB VRAM |
| DeepSeek-R1 Distill 70B | Q4_K_M | 43 GB | 42 GB | 81 GB | Larger than we track |
| DeepSeek-R1 0528 8B | Q4_K_M | 5.2 GB | 5.9 GB | 23 GB | 8 GB VRAM |
| SmolLM2 360M | Q4_K_M | 0.271 GB | 0.9 GB | 1.1 GB | 8 GB VRAM |
| SmolLM2 1.7B | Q4_K_M | 1.1 GB | 2.3 GB | 3 GB | 8 GB VRAM |
| OLMo 2 7B | Q4_K_M | 4.5 GB | 6.7 GB | 6.7 GB | 8 GB VRAM |
| OLMo 2 13B | Q4_K_M | 8.4 GB | 11 GB | 11 GB | 12 GB VRAM |
| TinyLlama 1.1B | Q4_K_M | 0.669 GB | 1.2 GB | 1.2 GB | 8 GB VRAM |
| Dolphin 3.0 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
How the memory is worked out
A local model needs three things to fit in your graphics card’s memory (VRAM):
- The model file. The weights are loaded as they are stored, so this is the download size its page shows. A smaller quantization (Q4 rather than Q8) is a smaller file, at some cost in quality.
- The context. For every token in the conversation the model keeps a key and a value for each layer and key/value head. We read the layers, heads and head width from the publisher’s own config file and count two bytes per number, which is how llama.cpp (and so Ollama and LM Studio) stores them by default. Double the context and this part doubles.
- An allowance of 0.5 GB for the program running the model. This one is our estimate, not a published figure.
Download pages give file sizes in decimal gigabytes (1,000,000,000 bytes), while graphics card and Mac memory is sold in binary ones (1,073,741,824 bytes). We show everything in the binary unit so it compares directly with your card, which is why a “5.2 GB” download reads as 4.8 GB here.
Graphics cards and Macs
A graphics card can give a model all of its VRAM, as the maker lists it. A Mac shares one pool of memory with macOS and your other apps, so we count 75% of it (also our estimate). Cards and Macs with the same memory get the same answer; how fast they run differs, but nobody publishes speed figures we could cite. See our methodology.
Frequently asked questions
How much VRAM do I need to run an LLM?
Roughly the size of the model file, plus memory for the context, plus a little for the program running it. Pick a model above to see each part. At a 4K context, the smallest graphics cards they fit on are: Qwen3 8B (Q4_K_M) on 8 GB, Qwen3 14B (Q4_K_M) on 12 GB, Qwen3 32B (Q4_K_M) on 24 GB.
Why does a longer context need more VRAM?
The model keeps a key and a value for every token in the conversation, in every layer. That memory grows in step with the context: double the context and it doubles. On some models a full-length context needs more memory than the model file itself.
What is quantization, and which one should I pick?
Quantization stores the model’s numbers with fewer bits, which makes the file, and the memory it needs, smaller, at some cost in quality. Q4 files are the smallest we list; Q8_0 keeps more detail but is a much bigger download. Where we list both, switch between them above to compare.
Can I run a model that needs more VRAM than my card has?
Often yes: Ollama and LM Studio can keep part of the model in your computer’s RAM. It works, but much more slowly, because that part runs on the processor. A Mac has no separate memory to spill into.
How accurate is this?
The file sizes and model designs come from the download pages and the publishers’ own config files, with the date we checked them. The allowance for the running program and the share of a Mac’s memory a model can use are our estimates, so a model shown as only just fitting may not.