How much VRAM does Qwen3 32B need?
At a 4K context, Qwen3 32B needs 20 GB of GPU memory with the Q4_K_M version and 34 GB with Q8_0. The smallest graphics card it fits on entirely has 24 GB of VRAM.
- Model file 19 GB listed as 20 GB
- Context 1 GB
- Allowance 0.5 GB
Graphics cards
- 8 GB VRAMGeForce RTX 5060 Ti 8 GBGeForce RTX 5060GeForce RTX 5050 +2 more 12 GB short
- 12 GB VRAMGeForce RTX 5070GeForce RTX 4070 TiGeForce RTX 4070 SUPER +1 more 8.1 GB short
- 16 GB VRAMGeForce RTX 5080GeForce RTX 5070 TiGeForce RTX 5060 Ti 16 GB +4 more 4.1 GB short
- 24 GB VRAMGeForce RTX 4090 3.9 GB spare
- 32 GB VRAMGeForce RTX 5090 12 GB spare
Macs
- 16 GB 12 GB usableMacBook Air (M5), 16 GBMac mini (M6), 16 GB 8.1 GB short
- 24 GB 18 GB usable
- 32 GB 24 GB usableMacBook Air (M5), 32 GBMac mini (M6), 32 GB 3.9 GB spare
- 48 GB 36 GB usableMac mini (M5 Pro), 48 GB 16 GB spare
- 64 GB 48 GB usableMac mini (M5 Pro), 64 GB 28 GB spare
Frequently asked questions
How much VRAM does Qwen3 32B need?
Q4_K_M (a 20 GB download) needs 20 GB at a 4K context and 29 GB at its full 40,960-token context; Q8_0 (a 35 GB download) needs 34 GB at a 4K context and 43 GB at its full 40,960-token context. That’s the model file, the memory its context needs and a 0.5 GB allowance for the program running it.
What’s the smallest GPU that can run Qwen3 32B?
A graphics card with 24 GB of VRAM, such as the GeForce RTX 4090, fits the Q4_K_M version at a 4K context with 3.9 GB to spare, and leaves room for up to 19,962 tokens of context. On a smaller card it can still run with part of the model in system RAM, but much more slowly.
Can Qwen3 32B run on a Mac?
Yes, on a Mac with 32 GB of unified memory or more, such as the MacBook Air (M5), 32 GB and Mac mini (M6), 32 GB. A Mac shares its memory with macOS and your apps, so we count about 24 GB of that 32 GB as usable.
How much memory does Qwen3 32B’s context take?
About 250 MB for every 1,000 tokens, worked out from Alibaba’s own config file. At its full 40,960-token context that adds 10 GB on top of the model file.
How fast will Qwen3 32B run?
Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show the memory it needs and where it fits.
| Model | Quantization | File | At 4K context | At max context | Smallest card |
|---|---|---|---|---|---|
| Qwen3 0.6B | Q4_K_M | 0.523 GB | 1.4 GB | 5.4 GB | 8 GB VRAM |
| Qwen3 1.7B | Q4_K_M | 1.4 GB | 2.2 GB | 6.2 GB | 8 GB VRAM |
| Qwen3 4B | Q4_K_M | 2.6 GB | 3.5 GB | 8.5 GB | 8 GB VRAM |
| Qwen3 8B | Q4_K_M | 5.2 GB | 5.9 GB | 11 GB | 8 GB VRAM |
| Q8_0 | 8.9 GB | 9.4 GB | 14 GB | 12 GB VRAM | |
| Qwen3 14B | Q4_K_M | 9.3 GB | 9.8 GB | 15 GB | 12 GB VRAM |
| Q8_0 | 16 GB | 16 GB | 22 GB | 24 GB VRAM | |
| Qwen3 32B | Q4_K_M | 20 GB | 20 GB | 29 GB | 24 GB VRAM |
| Q8_0 | 35 GB | 34 GB | 43 GB | Larger than we track | |
| Qwen2.5 0.5B | Q4_K_M | 0.398 GB | 0.9 GB | 1.2 GB | 8 GB VRAM |
| Qwen2.5 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Qwen2.5 72B | Q4_K_M | 47 GB | 46 GB | 54 GB | Larger than we track |
| Qwen2.5 Coder 1.5B | Q4_K_M | 0.986 GB | 1.5 GB | 2.3 GB | 8 GB VRAM |
| Qwen2.5 Coder 3B | Q4_K_M | 1.9 GB | 2.4 GB | 3.4 GB | 8 GB VRAM |
| Qwen2.5 Coder 7B | Q4_K_M | 4.7 GB | 5.1 GB | 6.6 GB | 8 GB VRAM |
| Qwen2.5 Coder 14B | Q4_K_M | 9 GB | 9.6 GB | 15 GB | 12 GB VRAM |
| Qwen2.5 Coder 32B | Q4_K_M | 20 GB | 20 GB | 27 GB | 24 GB VRAM |
| Phi-4 14B | Q4_K_M | 9.1 GB | 9.8 GB | 12 GB | 12 GB VRAM |
| Phi-4 mini 3.8B | Q4_K_M | 2.5 GB | 3.3 GB | 19 GB | 8 GB VRAM |
| Phi-3 mini 3.8B | Q4_K_M | 2.4 GB | 4.2 GB | 4.2 GB | 8 GB VRAM |
| Phi-3 medium 14B | Q4_K_M | 8.6 GB | 9.3 GB | 34 GB | 12 GB VRAM |
| Mistral 7B | Q4_K_M | 4.4 GB | 5.1 GB | 8.6 GB | 8 GB VRAM |
| Mistral Nemo 12B | Q4_K_M | 7.5 GB | 8.1 GB | 27 GB | 12 GB VRAM |
| Mistral Small 22B | Q4_K_M | 13 GB | 13 GB | 20 GB | 16 GB VRAM |
| Mistral Small 24B | Q4_K_M | 14 GB | 14 GB | 19 GB | 16 GB VRAM |
| DeepSeek-R1 Distill 1.5B | Q4_K_M | 1.1 GB | 1.6 GB | 5 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 7B | Q4_K_M | 4.7 GB | 5.1 GB | 12 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
| DeepSeek-R1 Distill 14B | Q4_K_M | 9 GB | 9.6 GB | 33 GB | 12 GB VRAM |
| DeepSeek-R1 Distill 32B | Q4_K_M | 20 GB | 20 GB | 51 GB | 24 GB VRAM |
| DeepSeek-R1 Distill 70B | Q4_K_M | 43 GB | 42 GB | 81 GB | Larger than we track |
| DeepSeek-R1 0528 8B | Q4_K_M | 5.2 GB | 5.9 GB | 23 GB | 8 GB VRAM |
| SmolLM2 360M | Q4_K_M | 0.271 GB | 0.9 GB | 1.1 GB | 8 GB VRAM |
| SmolLM2 1.7B | Q4_K_M | 1.1 GB | 2.3 GB | 3 GB | 8 GB VRAM |
| OLMo 2 7B | Q4_K_M | 4.5 GB | 6.7 GB | 6.7 GB | 8 GB VRAM |
| OLMo 2 13B | Q4_K_M | 8.4 GB | 11 GB | 11 GB | 12 GB VRAM |
| TinyLlama 1.1B | Q4_K_M | 0.669 GB | 1.2 GB | 1.2 GB | 8 GB VRAM |
| Dolphin 3.0 8B | Q4_K_M | 4.9 GB | 5.6 GB | 21 GB | 8 GB VRAM |
How the memory is worked out
A local model needs three things to fit in your graphics card’s memory (VRAM):
- The model file. The weights are loaded as they are stored, so this is the download size its page shows. A smaller quantization (Q4 rather than Q8) is a smaller file, at some cost in quality.
- The context. For every token in the conversation the model keeps a key and a value for each layer and key/value head. We read the layers, heads and head width from the publisher’s own config file and count two bytes per number, which is how llama.cpp (and so Ollama and LM Studio) stores them by default. Double the context and this part doubles.
- An allowance of 0.5 GB for the program running the model. This one is our estimate, not a published figure.
Download pages give file sizes in decimal gigabytes (1,000,000,000 bytes), while graphics card and Mac memory is sold in binary ones (1,073,741,824 bytes). We show everything in the binary unit so it compares directly with your card, which is why a “5.2 GB” download reads as 4.8 GB here.
Graphics cards and Macs
A graphics card can give a model all of its VRAM, as the maker lists it. A Mac shares one pool of memory with macOS and your other apps, so we count 75% of it (also our estimate). Cards and Macs with the same memory get the same answer; how fast they run differs, but nobody publishes speed figures we could cite. See our methodology.