How much VRAM does DeepSeek-R1 Distill 70B need?
At a 4K context, DeepSeek-R1 Distill 70B needs 42 GB of GPU memory with the Q4_K_M version. None of the graphics cards we track can hold it entirely.
- Model file 40 GB listed as 43 GB
- Context 1.3 GB
- Allowance 0.5 GB
Graphics cards
- 8 GB VRAMGeForce RTX 5060 Ti 8 GBGeForce RTX 5060GeForce RTX 5050 +2 more 34 GB short
- 12 GB VRAMGeForce RTX 5070GeForce RTX 4070 TiGeForce RTX 4070 SUPER +1 more 30 GB short
- 16 GB VRAMGeForce RTX 5080GeForce RTX 5070 TiGeForce RTX 5060 Ti 16 GB +4 more 26 GB short
- 24 GB VRAMGeForce RTX 4090 18 GB short
- 32 GB VRAMGeForce RTX 5090 9.8 GB short
Macs
- 16 GB 12 GB usableMacBook Air (M5), 16 GBMac mini (M6), 16 GB 30 GB short
- 24 GB 18 GB usable
- 32 GB 24 GB usableMacBook Air (M5), 32 GBMac mini (M6), 32 GB 18 GB short
- 48 GB 36 GB usableMac mini (M5 Pro), 48 GB 5.8 GB short
- 64 GB 48 GB usableMac mini (M5 Pro), 64 GB 6.2 GB spare
Frequently asked questions
How much VRAM does DeepSeek-R1 Distill 70B need?
Q4_K_M (a 43 GB download) needs 42 GB at a 4K context and 81 GB at its full 131,072-token context. That’s the model file, the memory its context needs and a 0.5 GB allowance for the program running it.
What’s the smallest GPU that can run DeepSeek-R1 Distill 70B?
None of the graphics cards we track have enough VRAM for the Q4_K_M version at a 4K context. It can still run with part of the model in system RAM, but much more slowly.
Can DeepSeek-R1 Distill 70B run on a Mac?
Yes, on a Mac with 64 GB of unified memory or more, such as the Mac mini (M5 Pro), 64 GB. A Mac shares its memory with macOS and your apps, so we count about 48 GB of that 64 GB as usable.
How much memory does DeepSeek-R1 Distill 70B’s context take?
About 313 MB for every 1,000 tokens, worked out from DeepSeek’s own config file. At its full 131,072-token context that adds 40 GB on top of the model file.
How fast will DeepSeek-R1 Distill 70B run?
Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show the memory it needs and where it fits.
| Model | Quantization | File | At 4K context | At max context | Smallest card |
|---|---|---|---|---|---|
| Qwen3 0.6B | Q4_K_M | 0.523 GB | 1.4 GB | 5.4 GB | 8 GB VRAM |
| Qwen3 1.7B | Q4_K_M | 1.4 GB | 2.2 GB | 6.2 GB | 8 GB VRAM |
| Qwen3 4B | Q4_K_M | 2.6 GB | 3.5 GB | 8.5 GB | 8 GB VRAM |
| Qwen3 8B | Q4_K_M | 5.2 GB | 5.9 GB | 11 GB | 8 GB VRAM |
| Q8_0 | 8.9 GB | 9.4 GB | 14 GB | 12 GB VRAM | |
| Qwen3 14B | Q4_K_M | 9.3 GB | 9.8 GB | 15 GB | 12 GB VRAM |
| Q8_0 | 16 GB | 16 GB | 22 GB | 24 GB VRAM | |
| Qwen3 32B | Q4_K_M | 20 GB | 20 GB | 29 GB | 24 GB VRAM |
| Q8_0 | 35 GB | 34 GB | 43 GB | Larger than we track | |
| Qwen2.5 0.5B | Q4_K_M | 0.398 GB | 0.9 GB | 1.2 GB | 8 GB VRAM |
| Qwen2.5 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Qwen2.5 72B | Q4_K_M | 47 GB | 46 GB | 54 GB | Larger than we track |
| Qwen2.5 Coder 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 Coder 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 Coder 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 Coder 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 Coder 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Phi-4 14B | Q4_K_M | 9.1 GB | 9.8 GB | 12 GB | 12 GB VRAM |
| Phi-4 mini 3.8B | Q4_K_M | 2.5 GB | 3.3 GB | 19 GB | 8 GB VRAM |
| Phi-3 mini 3.8B | Q4_K_M | 2.4 GB | 4.2 GB | 4.2 GB | 8 GB VRAM |
| Phi-3 medium 14B | Q4_K_M | 8.6 GB | 9.3 GB | 34 GB | 12 GB VRAM |
| Mistral 7B | Q4_K_M | 4.4 GB | 5.1 GB | 8.6 GB | 8 GB VRAM |
| Mistral Nemo 12B | Q4_K_M | 7.5 GB | 8.1 GB | 27 GB | 12 GB VRAM |
| Mistral Small 22B | Q4_K_M | 13 GB | 13 GB | 20 GB | 16 GB VRAM |
| Mistral Small 24B | Q4_K_M | 14 GB | 14 GB | 19 GB | 16 GB VRAM |
| DeepSeek-R1 Distill 1.5B | Q4_K_M | 1.1 GB | 1.6 GB | 5 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 7B | Q4_K_M | 4.7 GB | 5.1 GB | 12 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 14B | Q4_K_M | 9 GB | 9.6 GB | 33 GB | 12 GB VRAM |
| DeepSeek-R1 Distill 32B | Q4_K_M | 20 GB | 20 GB | 51 GB | 24 GB VRAM |
| DeepSeek-R1 Distill 70B | Q4_K_M | 43 GB | 42 GB | 81 GB | Larger than we track |
| DeepSeek-R1 0528 8B | Q4_K_M | 5.2 GB | 5.9 GB | 23 GB | 8 GB VRAM |
| SmolLM2 360M | Q4_K_M | 0.271 GB | 0.9 GB | 1.1 GB | 8 GB VRAM |
| SmolLM2 1.7B | Q4_K_M | 1.1 GB | 2.3 GB | 3 GB | 8 GB VRAM |
| OLMo 2 7B | Q4_K_M | 4.5 GB | 6.7 GB | 6.7 GB | 8 GB VRAM |
| OLMo 2 13B | Q4_K_M | 8.4 GB | 11 GB | 11 GB | 12 GB VRAM |
| TinyLlama 1.1B | Q4_K_M | 0.669 GB | 1.2 GB | 1.2 GB | 8 GB VRAM |
| Dolphin 3.0 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
How the memory is worked out
A local model needs three things to fit in your graphics card’s memory (VRAM):
- The model file. The weights are loaded as they are stored, so this is the download size its page shows. A smaller quantization (Q4 rather than Q8) is a smaller file, at some cost in quality.
- The context. For every token in the conversation the model keeps a key and a value for each layer and key/value head. We read the layers, heads and head width from the publisher’s own config file and count two bytes per number, which is how llama.cpp (and so Ollama and LM Studio) stores them by default. Double the context and this part doubles.
- An allowance of 0.5 GB for the program running the model. This one is our estimate, not a published figure.
Download pages give file sizes in decimal gigabytes (1,000,000,000 bytes), while graphics card and Mac memory is sold in binary ones (1,073,741,824 bytes). We show everything in the binary unit so it compares directly with your card, which is why a “5.2 GB” download reads as 4.8 GB here.
Graphics cards and Macs
A graphics card can give a model all of its VRAM, as the maker lists it. A Mac shares one pool of memory with macOS and your other apps, so we count 75% of it (also our estimate). Cards and Macs with the same memory get the same answer; how fast they run differs, but nobody publishes speed figures we could cite. See our methodology.