Samprix

What AI can the GeForce RTX 4080 SUPER run?

The GeForce RTX 4080 SUPER can give a local model 16 GB of memory. At a 4K context, 34 of the 42 models we track fit entirely on it; the largest is Mistral Small 24B.

NVIDIA Specs last checked: how
Your computer
Context length
System RAM
Memory a model can use 16 GB 16 GB of VRAM · source

Models on this machine

  • Model file
  • Context
  • Allowance
  • Your GPU memory

Fits

34
  • Mistral Small 24B
    Mistral AI · Q4_K_M
    Needs 14 GB 1.8 GB spare
    Context up to
    16,131
  • Mistral Small 22B
    Mistral AI · Q4_K_M
    Needs 13 GB 2.5 GB spare
    Context up to
    15,882
  • Qwen3 14B
    Alibaba · Q4_K_M
    Needs 9.8 GB 6.2 GB spare
    Context up to
    40,960
  • Phi-4 14B
    Microsoft · Q4_K_M
    Needs 9.8 GB 6.2 GB spare
    Context up to
    16,384
  • Qwen2.5 14B
    Alibaba · Q4_K_M
    Needs 9.6 GB 6.4 GB spare
    Context up to
    32,768
  • DeepSeek-R1 Distill 14B
    DeepSeek · Q4_K_M
    Needs 9.6 GB 6.4 GB spare
    Context up to
    38,874
  • Qwen2.5 Coder 14B
    Alibaba · Q4_K_M
    Needs 9.6 GB 6.4 GB spare
    Context up to
    32,768
  • Qwen3 8B
    Alibaba · Q8_0
    Needs 9.4 GB 6.6 GB spare
    Context up to
    40,960
  • Phi-3 medium 14B
    Microsoft · Q4_K_M
    Needs 9.3 GB 6.7 GB spare
    Context up to
    39,272
  • OLMo 2 13B
    Allen Institute for AI · Q4_K_M
    Needs 11 GB 4.6 GB spare
    Context up to
    4,096
  • Mistral Nemo 12B
    Mistral AI · Q4_K_M
    Needs 8.1 GB 7.9 GB spare
    Context up to
    55,804
  • Qwen3 8B
    Alibaba · Q4_K_M
    Needs 5.9 GB 10 GB spare
    Context up to
    40,960
  • DeepSeek-R1 0528 8B
    DeepSeek · Q4_K_M
    Needs 5.9 GB 10 GB spare
    Context up to
    77,602
  • DeepSeek-R1 Distill 8B
    DeepSeek · Q4_K_M
    Needs 5.6 GB 10 GB spare
    Context up to
    89,591
  • Dolphin 3.0 8B
    Cognitive Computations · Q4_K_M
    Needs 5.6 GB 10 GB spare
    Context up to
    89,591
  • Qwen2.5 7B
    Alibaba · Q4_K_M
    Needs 5.1 GB 11 GB spare
    Context up to
    32,768
  • Qwen2.5 Coder 7B
    Alibaba · Q4_K_M
    Needs 5.1 GB 11 GB spare
    Context up to
    32,768
  • DeepSeek-R1 Distill 7B
    DeepSeek · Q4_K_M
    Needs 5.1 GB 11 GB spare
    Context up to
    131,072
  • OLMo 2 7B
    Allen Institute for AI · Q4_K_M
    Needs 6.7 GB 9.3 GB spare
    Context up to
    4,096
  • Mistral 7B
    Mistral AI · Q4_K_M
    Needs 5.1 GB 11 GB spare
    Context up to
    32,768
  • Qwen3 4B
    Alibaba · Q4_K_M
    Needs 3.5 GB 13 GB spare
    Context up to
    40,960
  • Phi-4 mini 3.8B
    Microsoft · Q4_K_M
    Needs 3.3 GB 13 GB spare
    Context up to
    107,902
  • Phi-3 mini 3.8B
    Microsoft · Q4_K_M
    Needs 4.2 GB 12 GB spare
    Context up to
    4,096
  • Qwen2.5 3B
    Alibaba · Q4_K_M
    Needs 2.4 GB 14 GB spare
    Context up to
    32,768
  • Qwen2.5 Coder 3B
    Alibaba · Q4_K_M
    Needs 2.4 GB 14 GB spare
    Context up to
    32,768
  • Qwen3 1.7B
    Alibaba · Q4_K_M
    Needs 2.2 GB 14 GB spare
    Context up to
    40,960
  • DeepSeek-R1 Distill 1.5B
    DeepSeek · Q4_K_M
    Needs 1.6 GB 14 GB spare
    Context up to
    131,072
  • SmolLM2 1.7B
    Hugging Face · Q4_K_M
    Needs 2.3 GB 14 GB spare
    Context up to
    8,192
  • Qwen2.5 1.5B
    Alibaba · Q4_K_M
    Needs 1.5 GB 14 GB spare
    Context up to
    32,768
  • Qwen2.5 Coder 1.5B
    Alibaba · Q4_K_M
    Needs 1.5 GB 14 GB spare
    Context up to
    32,768
  • TinyLlama 1.1B
    TinyLlama · Q4_K_M
    Needs 1.2 GB 15 GB spare
    Context up to
    2,048
  • Qwen3 0.6B
    Alibaba · Q4_K_M
    Needs 1.4 GB 15 GB spare
    Context up to
    40,960
  • Qwen2.5 0.5B
    Alibaba · Q4_K_M
    Needs 0.9 GB 15 GB spare
    Context up to
    32,768
  • SmolLM2 360M
    Hugging Face · Q4_K_M
    Needs 0.9 GB 15 GB spare
    Context up to
    8,192

Won’t fit

8
Memory needed = model file + context + a 0.5 GB allowance (our estimate). See methodology.

Frequently asked questions

What’s the biggest AI model the GeForce RTX 4080 SUPER can run?

Mistral Small 24B (Q4_K_M, a 14 GB download) is the largest of the 42 models we track that fits entirely in the 16 GB the GeForce RTX 4080 SUPER can give a model at a 4K context, with 1.8 GB to spare and room for up to 16,131 tokens of context.

Can the GeForce RTX 4080 SUPER run Qwen3 32B?

Not entirely in its own memory. Qwen3 32B (Q4_K_M) needs 20 GB at a 4K context, and the GeForce RTX 4080 SUPER can give a model 16 GB. With enough system RAM the rest can spill over into it, which works but is much slower; pick your RAM above to check.

How much memory can the GeForce RTX 4080 SUPER give an AI model?

All 16 GB of its VRAM, per NVIDIA GeForce graphics card comparison (checked Oct 2, 2026). The model file, its context and the program running it all have to fit in that.

Which other graphics cards or Macs run the same models?

GeForce RTX 5080, GeForce RTX 5070 Ti, GeForce RTX 5060 Ti 16 GB, GeForce RTX 4080, GeForce RTX 4070 Ti SUPER and GeForce RTX 4060 Ti 16 GB have the same 16 GB of VRAM, so the same models fit. How fast they run differs, but makers don’t publish speed figures we could cite.

How fast will local AI run on the GeForce RTX 4080 SUPER?

Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show whether a model fits and how much context it leaves room for.

Other graphics cards and Macs

Did this tool do what you needed?

© 2026 Samprix. Not affiliated with OpenAI, Anthropic or Google.