What AI can the GeForce RTX 4090 run?
The GeForce RTX 4090 can give a local model 24 GB of memory. At a 4K context, 39 of the 42 models we track fit entirely on it; the largest is Qwen3 32B.
- Model file
- Context
- Allowance
- Your GPU memory
Fits
39- Qwen3 32BAlibaba · Q4_K_MNeeds 20 GB 3.9 GB spareContext up to19,962
- Qwen2.5 32BAlibaba · Q4_K_MNeeds 20 GB 3.9 GB spareContext up to19,962
- DeepSeek-R1 Distill 32BDeepSeek · Q4_K_MNeeds 20 GB 3.9 GB spareContext up to19,962
- Qwen2.5 Coder 32BAlibaba · Q4_K_MNeeds 20 GB 3.9 GB spareContext up to19,962
- Qwen3 14BAlibaba · Q8_0Needs 16 GB 8 GB spareContext up to40,960
- Mistral Small 24BMistral AI · Q4_K_MNeeds 14 GB 9.8 GB spareContext up to32,768
- Mistral Small 22BMistral AI · Q4_K_MNeeds 13 GB 11 GB spareContext up to32,768
- Qwen3 14BAlibaba · Q4_K_MNeeds 9.8 GB 14 GB spareContext up to40,960
- Phi-4 14BMicrosoft · Q4_K_MNeeds 9.8 GB 14 GB spareContext up to16,384
- Qwen2.5 14BAlibaba · Q4_K_MNeeds 9.6 GB 14 GB spareContext up to32,768
- DeepSeek-R1 Distill 14BDeepSeek · Q4_K_MNeeds 9.6 GB 14 GB spareContext up to82,564
- Qwen2.5 Coder 14BAlibaba · Q4_K_MNeeds 9.6 GB 14 GB spareContext up to32,768
- Qwen3 8BAlibaba · Q8_0Needs 9.4 GB 15 GB spareContext up to40,960
- Phi-3 medium 14BMicrosoft · Q4_K_MNeeds 9.3 GB 15 GB spareContext up to81,215
- OLMo 2 13BAllen Institute for AI · Q4_K_MNeeds 11 GB 13 GB spareContext up to4,096
- Mistral Nemo 12BMistral AI · Q4_K_MNeeds 8.1 GB 16 GB spareContext up to108,233
- Qwen3 8BAlibaba · Q4_K_MNeeds 5.9 GB 18 GB spareContext up to40,960
- DeepSeek-R1 0528 8BDeepSeek · Q4_K_MNeeds 5.9 GB 18 GB spareContext up to131,072
- DeepSeek-R1 Distill 8BDeepSeek · Q4_K_MNeeds 5.6 GB 18 GB spareContext up to131,072
- Dolphin 3.0 8BCognitive Computations · Q4_K_MNeeds 5.6 GB 18 GB spareContext up to131,072
- Qwen2.5 7BAlibaba · Q4_K_MNeeds 5.1 GB 19 GB spareContext up to32,768
- Qwen2.5 Coder 7BAlibaba · Q4_K_MNeeds 5.1 GB 19 GB spareContext up to32,768
- DeepSeek-R1 Distill 7BDeepSeek · Q4_K_MNeeds 5.1 GB 19 GB spareContext up to131,072
- OLMo 2 7BAllen Institute for AI · Q4_K_MNeeds 6.7 GB 17 GB spareContext up to4,096
- Mistral 7BMistral AI · Q4_K_MNeeds 5.1 GB 19 GB spareContext up to32,768
- Qwen3 4BAlibaba · Q4_K_MNeeds 3.5 GB 21 GB spareContext up to40,960
- Phi-4 mini 3.8BMicrosoft · Q4_K_MNeeds 3.3 GB 21 GB spareContext up to131,072
- Phi-3 mini 3.8BMicrosoft · Q4_K_MNeeds 4.2 GB 20 GB spareContext up to4,096
- Qwen2.5 3BAlibaba · Q4_K_MNeeds 2.4 GB 22 GB spareContext up to32,768
- Qwen2.5 Coder 3BAlibaba · Q4_K_MNeeds 2.4 GB 22 GB spareContext up to32,768
- Qwen3 1.7BAlibaba · Q4_K_MNeeds 2.2 GB 22 GB spareContext up to40,960
- DeepSeek-R1 Distill 1.5BDeepSeek · Q4_K_MNeeds 1.6 GB 22 GB spareContext up to131,072
- SmolLM2 1.7BHugging Face · Q4_K_MNeeds 2.3 GB 22 GB spareContext up to8,192
- Qwen2.5 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 22 GB spareContext up to32,768
- Qwen2.5 Coder 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 22 GB spareContext up to32,768
- TinyLlama 1.1BTinyLlama · Q4_K_MNeeds 1.2 GB 23 GB spareContext up to2,048
- Qwen3 0.6BAlibaba · Q4_K_MNeeds 1.4 GB 23 GB spareContext up to40,960
- Qwen2.5 0.5BAlibaba · Q4_K_MNeeds 0.9 GB 23 GB spareContext up to32,768
- SmolLM2 360MHugging Face · Q4_K_MNeeds 0.9 GB 23 GB spareContext up to8,192
Won’t fit
3- Qwen2.5 72BAlibaba · Q4_K_MNeeds 46 GB 22 GB short
- DeepSeek-R1 Distill 70BDeepSeek · Q4_K_MNeeds 42 GB 18 GB short
- Qwen3 32BAlibaba · Q8_0Needs 34 GB 10 GB short
Frequently asked questions
What’s the biggest AI model the GeForce RTX 4090 can run?
Qwen3 32B (Q4_K_M, a 20 GB download) is the largest of the 42 models we track that fits entirely in the 24 GB the GeForce RTX 4090 can give a model at a 4K context, with 3.9 GB to spare and room for up to 19,962 tokens of context.
Can the GeForce RTX 4090 run DeepSeek-R1 Distill 70B?
Not entirely in its own memory. DeepSeek-R1 Distill 70B (Q4_K_M) needs 42 GB at a 4K context, and the GeForce RTX 4090 can give a model 24 GB. With enough system RAM the rest can spill over into it, which works but is much slower; pick your RAM above to check.
How much memory can the GeForce RTX 4090 give an AI model?
All 24 GB of its VRAM, per NVIDIA GeForce graphics card comparison (checked Oct 2, 2026). The model file, its context and the program running it all have to fit in that.
How fast will local AI run on the GeForce RTX 4090?
Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show whether a model fits and how much context it leaves room for.
Other graphics cards and Macs
NVIDIA
- GeForce RTX 5090
- GeForce RTX 5080
- GeForce RTX 5070 Ti
- GeForce RTX 5070
- GeForce RTX 5060 Ti 16 GB
- GeForce RTX 5060 Ti 8 GB
- GeForce RTX 5060
- GeForce RTX 5050
- GeForce RTX 4090
- GeForce RTX 4080 SUPER
- GeForce RTX 4080
- GeForce RTX 4070 Ti SUPER
- GeForce RTX 4070 Ti
- GeForce RTX 4070 SUPER
- GeForce RTX 4070
- GeForce RTX 4060 Ti 16 GB
- GeForce RTX 4060 Ti 8 GB
- GeForce RTX 4060
Did this tool do what you needed?
Related tools
Set up local AI on Windows or Mac, step by step.
What running a model on your own hardware costs, compared with the API.
See how many messages a day it takes before a ChatGPT, Claude or Gemini plan beats paying per token.