What AI can the GeForce RTX 4070 run?
The GeForce RTX 4070 can give a local model 12 GB of memory. At a 4K context, 32 of the 42 models we track fit entirely on it; the largest is Qwen3 14B.
- Model file
- Context
- Allowance
- Your GPU memory
Fits
32- Qwen3 14BAlibaba · Q4_K_MNeeds 9.8 GB 2.2 GB spareContext up to18,603
- Phi-4 14BMicrosoft · Q4_K_MNeeds 9.8 GB 2.2 GB spareContext up to15,859
- Qwen2.5 14BAlibaba · Q4_K_MNeeds 9.6 GB 2.4 GB spareContext up to17,028
- DeepSeek-R1 Distill 14BDeepSeek · Q4_K_MNeeds 9.6 GB 2.4 GB spareContext up to17,028
- Qwen2.5 Coder 14BAlibaba · Q4_K_MNeeds 9.6 GB 2.4 GB spareContext up to17,028
- Qwen3 8BAlibaba · Q8_0Needs 9.4 GB 2.6 GB spareContext up to23,383
- Phi-3 medium 14BMicrosoft · Q4_K_MNeeds 9.3 GB 2.7 GB spareContext up to18,300
- OLMo 2 13BAllen Institute for AI · Q4_K_MNeeds 11 GB 0.6 GB spareContext up to4,096
- Mistral Nemo 12BMistral AI · Q4_K_MNeeds 8.1 GB 3.9 GB spareContext up to29,590
- Qwen3 8BAlibaba · Q4_K_MNeeds 5.9 GB 6.1 GB spareContext up to40,960
- DeepSeek-R1 0528 8BDeepSeek · Q4_K_MNeeds 5.9 GB 6.1 GB spareContext up to48,475
- DeepSeek-R1 Distill 8BDeepSeek · Q4_K_MNeeds 5.6 GB 6.4 GB spareContext up to56,823
- Dolphin 3.0 8BCognitive Computations · Q4_K_MNeeds 5.6 GB 6.4 GB spareContext up to56,823
- Qwen2.5 7BAlibaba · Q4_K_MNeeds 5.1 GB 6.9 GB spareContext up to32,768
- Qwen2.5 Coder 7BAlibaba · Q4_K_MNeeds 5.1 GB 6.9 GB spareContext up to32,768
- DeepSeek-R1 Distill 7BDeepSeek · Q4_K_MNeeds 5.1 GB 6.9 GB spareContext up to131,072
- OLMo 2 7BAllen Institute for AI · Q4_K_MNeeds 6.7 GB 5.3 GB spareContext up to4,096
- Mistral 7BMistral AI · Q4_K_MNeeds 5.1 GB 6.9 GB spareContext up to32,768
- Qwen3 4BAlibaba · Q4_K_MNeeds 3.5 GB 8.5 GB spareContext up to40,960
- Phi-4 mini 3.8BMicrosoft · Q4_K_MNeeds 3.3 GB 8.7 GB spareContext up to75,134
- Phi-3 mini 3.8BMicrosoft · Q4_K_MNeeds 4.2 GB 7.8 GB spareContext up to4,096
- Qwen2.5 3BAlibaba · Q4_K_MNeeds 2.4 GB 9.6 GB spareContext up to32,768
- Qwen2.5 Coder 3BAlibaba · Q4_K_MNeeds 2.4 GB 9.6 GB spareContext up to32,768
- Qwen3 1.7BAlibaba · Q4_K_MNeeds 2.2 GB 9.8 GB spareContext up to40,960
- DeepSeek-R1 Distill 1.5BDeepSeek · Q4_K_MNeeds 1.6 GB 10 GB spareContext up to131,072
- SmolLM2 1.7BHugging Face · Q4_K_MNeeds 2.3 GB 9.7 GB spareContext up to8,192
- Qwen2.5 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 10 GB spareContext up to32,768
- Qwen2.5 Coder 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 10 GB spareContext up to32,768
- TinyLlama 1.1BTinyLlama · Q4_K_MNeeds 1.2 GB 11 GB spareContext up to2,048
- Qwen3 0.6BAlibaba · Q4_K_MNeeds 1.4 GB 11 GB spareContext up to40,960
- Qwen2.5 0.5BAlibaba · Q4_K_MNeeds 0.9 GB 11 GB spareContext up to32,768
- SmolLM2 360MHugging Face · Q4_K_MNeeds 0.9 GB 11 GB spareContext up to8,192
Won’t fit
10- Qwen2.5 72BAlibaba · Q4_K_MNeeds 46 GB 34 GB short
- DeepSeek-R1 Distill 70BDeepSeek · Q4_K_MNeeds 42 GB 30 GB short
- Qwen3 32BAlibaba · Q8_0Needs 34 GB 22 GB short
- Qwen3 32BAlibaba · Q4_K_MNeeds 20 GB 8.1 GB short
- Qwen2.5 32BAlibaba · Q4_K_MNeeds 20 GB 8.1 GB short
- DeepSeek-R1 Distill 32BDeepSeek · Q4_K_MNeeds 20 GB 8.1 GB short
- Qwen2.5 Coder 32BAlibaba · Q4_K_MNeeds 20 GB 8.1 GB short
- Qwen3 14BAlibaba · Q8_0Needs 16 GB 4 GB short
- Mistral Small 24BMistral AI · Q4_K_MNeeds 14 GB 2.2 GB short
- Mistral Small 22BMistral AI · Q4_K_MNeeds 13 GB 1.5 GB short
Frequently asked questions
What’s the biggest AI model the GeForce RTX 4070 can run?
Qwen3 14B (Q4_K_M, a 9.3 GB download) is the largest of the 42 models we track that fits entirely in the 12 GB the GeForce RTX 4070 can give a model at a 4K context, with 2.2 GB to spare and room for up to 18,603 tokens of context.
Can the GeForce RTX 4070 run Mistral Small 22B?
Not entirely in its own memory. Mistral Small 22B (Q4_K_M) needs 13 GB at a 4K context, and the GeForce RTX 4070 can give a model 12 GB. With enough system RAM the rest can spill over into it, which works but is much slower; pick your RAM above to check.
How much memory can the GeForce RTX 4070 give an AI model?
All 12 GB of its VRAM, per NVIDIA GeForce graphics card comparison (checked Oct 2, 2026). The model file, its context and the program running it all have to fit in that.
Which other graphics cards or Macs run the same models?
GeForce RTX 5070, GeForce RTX 4070 Ti and GeForce RTX 4070 SUPER have the same 12 GB of VRAM, so the same models fit. How fast they run differs, but makers don’t publish speed figures we could cite.
How fast will local AI run on the GeForce RTX 4070?
Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show whether a model fits and how much context it leaves room for.
Other graphics cards and Macs
NVIDIA
- GeForce RTX 5090
- GeForce RTX 5080
- GeForce RTX 5070 Ti
- GeForce RTX 5070
- GeForce RTX 5060 Ti 16 GB
- GeForce RTX 5060 Ti 8 GB
- GeForce RTX 5060
- GeForce RTX 5050
- GeForce RTX 4090
- GeForce RTX 4080 SUPER
- GeForce RTX 4080
- GeForce RTX 4070 Ti SUPER
- GeForce RTX 4070 Ti
- GeForce RTX 4070 SUPER
- GeForce RTX 4070
- GeForce RTX 4060 Ti 16 GB
- GeForce RTX 4060 Ti 8 GB
- GeForce RTX 4060
Did this tool do what you needed?
Related tools
Set up local AI on Windows or Mac, step by step.
What running a model on your own hardware costs, compared with the API.
See how many messages a day it takes before a ChatGPT, Claude or Gemini plan beats paying per token.