What AI can the GeForce RTX 5080 run?
The GeForce RTX 5080 can give a local model 16 GB of memory. At a 4K context, 34 of the 42 models we track fit entirely on it; the largest is Mistral Small 24B.
- Model file
- Context
- Allowance
- Your GPU memory
Fits
34- Mistral Small 24BMistral AI · Q4_K_MNeeds 14 GB 1.8 GB spareContext up to16,131
- Mistral Small 22BMistral AI · Q4_K_MNeeds 13 GB 2.5 GB spareContext up to15,882
- Qwen3 14BAlibaba · Q4_K_MNeeds 9.8 GB 6.2 GB spareContext up to40,960
- Phi-4 14BMicrosoft · Q4_K_MNeeds 9.8 GB 6.2 GB spareContext up to16,384
- Qwen2.5 14BAlibaba · Q4_K_MNeeds 9.6 GB 6.4 GB spareContext up to32,768
- DeepSeek-R1 Distill 14BDeepSeek · Q4_K_MNeeds 9.6 GB 6.4 GB spareContext up to38,874
- Qwen2.5 Coder 14BAlibaba · Q4_K_MNeeds 9.6 GB 6.4 GB spareContext up to32,768
- Qwen3 8BAlibaba · Q8_0Needs 9.4 GB 6.6 GB spareContext up to40,960
- Phi-3 medium 14BMicrosoft · Q4_K_MNeeds 9.3 GB 6.7 GB spareContext up to39,272
- OLMo 2 13BAllen Institute for AI · Q4_K_MNeeds 11 GB 4.6 GB spareContext up to4,096
- Mistral Nemo 12BMistral AI · Q4_K_MNeeds 8.1 GB 7.9 GB spareContext up to55,804
- Qwen3 8BAlibaba · Q4_K_MNeeds 5.9 GB 10 GB spareContext up to40,960
- DeepSeek-R1 0528 8BDeepSeek · Q4_K_MNeeds 5.9 GB 10 GB spareContext up to77,602
- DeepSeek-R1 Distill 8BDeepSeek · Q4_K_MNeeds 5.6 GB 10 GB spareContext up to89,591
- Dolphin 3.0 8BCognitive Computations · Q4_K_MNeeds 5.6 GB 10 GB spareContext up to89,591
- Qwen2.5 7BAlibaba · Q4_K_MNeeds 5.1 GB 11 GB spareContext up to32,768
- Qwen2.5 Coder 7BAlibaba · Q4_K_MNeeds 5.1 GB 11 GB spareContext up to32,768
- DeepSeek-R1 Distill 7BDeepSeek · Q4_K_MNeeds 5.1 GB 11 GB spareContext up to131,072
- OLMo 2 7BAllen Institute for AI · Q4_K_MNeeds 6.7 GB 9.3 GB spareContext up to4,096
- Mistral 7BMistral AI · Q4_K_MNeeds 5.1 GB 11 GB spareContext up to32,768
- Qwen3 4BAlibaba · Q4_K_MNeeds 3.5 GB 13 GB spareContext up to40,960
- Phi-4 mini 3.8BMicrosoft · Q4_K_MNeeds 3.3 GB 13 GB spareContext up to107,902
- Phi-3 mini 3.8BMicrosoft · Q4_K_MNeeds 4.2 GB 12 GB spareContext up to4,096
- Qwen2.5 3BAlibaba · Q4_K_MNeeds 2.4 GB 14 GB spareContext up to32,768
- Qwen2.5 Coder 3BAlibaba · Q4_K_MNeeds 2.4 GB 14 GB spareContext up to32,768
- Qwen3 1.7BAlibaba · Q4_K_MNeeds 2.2 GB 14 GB spareContext up to40,960
- DeepSeek-R1 Distill 1.5BDeepSeek · Q4_K_MNeeds 1.6 GB 14 GB spareContext up to131,072
- SmolLM2 1.7BHugging Face · Q4_K_MNeeds 2.3 GB 14 GB spareContext up to8,192
- Qwen2.5 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 14 GB spareContext up to32,768
- Qwen2.5 Coder 1.5BAlibaba · Q4_K_MNeeds 1.5 GB 14 GB spareContext up to32,768
- TinyLlama 1.1BTinyLlama · Q4_K_MNeeds 1.2 GB 15 GB spareContext up to2,048
- Qwen3 0.6BAlibaba · Q4_K_MNeeds 1.4 GB 15 GB spareContext up to40,960
- Qwen2.5 0.5BAlibaba · Q4_K_MNeeds 0.9 GB 15 GB spareContext up to32,768
- SmolLM2 360MHugging Face · Q4_K_MNeeds 0.9 GB 15 GB spareContext up to8,192
Won’t fit
8- Qwen2.5 72BAlibaba · Q4_K_MNeeds 46 GB 30 GB short
- DeepSeek-R1 Distill 70BDeepSeek · Q4_K_MNeeds 42 GB 26 GB short
- Qwen3 32BAlibaba · Q8_0Needs 34 GB 18 GB short
- Qwen3 32BAlibaba · Q4_K_MNeeds 20 GB 4.1 GB short
- Qwen2.5 32BAlibaba · Q4_K_MNeeds 20 GB 4.1 GB short
- DeepSeek-R1 Distill 32BDeepSeek · Q4_K_MNeeds 20 GB 4.1 GB short
- Qwen2.5 Coder 32BAlibaba · Q4_K_MNeeds 20 GB 4.1 GB short
- Qwen3 14BAlibaba · Q8_0Needs 16 GB 0 GB shortContext up to3,924
Frequently asked questions
What’s the biggest AI model the GeForce RTX 5080 can run?
Mistral Small 24B (Q4_K_M, a 14 GB download) is the largest of the 42 models we track that fits entirely in the 16 GB the GeForce RTX 5080 can give a model at a 4K context, with 1.8 GB to spare and room for up to 16,131 tokens of context.
Can the GeForce RTX 5080 run Qwen3 32B?
Not entirely in its own memory. Qwen3 32B (Q4_K_M) needs 20 GB at a 4K context, and the GeForce RTX 5080 can give a model 16 GB. With enough system RAM the rest can spill over into it, which works but is much slower; pick your RAM above to check.
How much memory can the GeForce RTX 5080 give an AI model?
All 16 GB of its VRAM, per NVIDIA GeForce graphics card comparison (checked Oct 2, 2026). The model file, its context and the program running it all have to fit in that.
Which other graphics cards or Macs run the same models?
GeForce RTX 5070 Ti, GeForce RTX 5060 Ti 16 GB, GeForce RTX 4080 SUPER, GeForce RTX 4080, GeForce RTX 4070 Ti SUPER and GeForce RTX 4060 Ti 16 GB have the same 16 GB of VRAM, so the same models fit. How fast they run differs, but makers don’t publish speed figures we could cite.
How fast will local AI run on the GeForce RTX 5080?
Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show whether a model fits and how much context it leaves room for.
Other graphics cards and Macs
NVIDIA
- GeForce RTX 5090
- GeForce RTX 5080
- GeForce RTX 5070 Ti
- GeForce RTX 5070
- GeForce RTX 5060 Ti 16 GB
- GeForce RTX 5060 Ti 8 GB
- GeForce RTX 5060
- GeForce RTX 5050
- GeForce RTX 4090
- GeForce RTX 4080 SUPER
- GeForce RTX 4080
- GeForce RTX 4070 Ti SUPER
- GeForce RTX 4070 Ti
- GeForce RTX 4070 SUPER
- GeForce RTX 4070
- GeForce RTX 4060 Ti 16 GB
- GeForce RTX 4060 Ti 8 GB
- GeForce RTX 4060
Did this tool do what you needed?
Related tools
Set up local AI on Windows or Mac, step by step.
What running a model on your own hardware costs, compared with the API.
See how many messages a day it takes before a ChatGPT, Claude or Gemini plan beats paying per token.