Samprix

How much VRAM does Qwen2.5 72B need?

At a 4K context, Qwen2.5 72B needs 46 GB of GPU memory with the Q4_K_M version. None of the graphics cards we track can hold it entirely.

Alibaba Specs last checked: how
Model
Quantization
Context length
Memory needed 46 GB Qwen2.5 72B · Q4_K_M · 4,096 tokens
  • Model file 44 GB listed as 47 GB
  • Context 1.3 GB
  • Allowance 0.5 GB
File size from Ollama library: qwen2.5; context memory from the publisher’s config, checked Oct 2, 2026.

Memory by context length

+313 MB per 1K tokens

Where it fits at 4K context

Memory available

Graphics cards

Macs

A Mac gives a model about 75% of its unified memory (our estimate). A card that’s short can still run it with part in system RAM, much more slowly. Check your machine

Frequently asked questions

How much VRAM does Qwen2.5 72B need?

Q4_K_M (a 47 GB download) needs 46 GB at a 4K context and 54 GB at its full 32,768-token context. That’s the model file, the memory its context needs and a 0.5 GB allowance for the program running it.

What’s the smallest GPU that can run Qwen2.5 72B?

None of the graphics cards we track have enough VRAM for the Q4_K_M version at a 4K context. It can still run with part of the model in system RAM, but much more slowly.

Can Qwen2.5 72B run on a Mac?

Yes, on a Mac with 64 GB of unified memory or more, such as the Mac mini (M5 Pro), 64 GB. A Mac shares its memory with macOS and your apps, so we count about 48 GB of that 64 GB as usable.

How much memory does Qwen2.5 72B’s context take?

About 313 MB for every 1,000 tokens, worked out from Alibaba’s own config file. At its full 32,768-token context that adds 10 GB on top of the model file.

How fast will Qwen2.5 72B run?

Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show the memory it needs and where it fits.

Did this tool do what you needed?

© 2026 Samprix. Not affiliated with OpenAI, Anthropic or Google.