Samprix

How much VRAM does Mistral Small 22B need?

At a 4K context, Mistral Small 22B needs 13 GB of GPU memory with the Q4_K_M version. The smallest graphics card it fits on entirely has 16 GB of VRAM.

Mistral AI Specs last checked: how
Model
Quantization
Context length
Memory needed 13 GB Mistral Small 22B · Q4_K_M · 4,096 tokens
  • Model file 12 GB listed as 13 GB
  • Context 0.9 GB
  • Allowance 0.5 GB
File size from Ollama library: mistral-small; context memory from the publisher’s config, checked Oct 2, 2026.

Memory by context length

+219 MB per 1K tokens

Where it fits at 4K context

Memory available

Graphics cards

Macs

A Mac gives a model about 75% of its unified memory (our estimate). A card that’s short can still run it with part in system RAM, much more slowly. Check your machine

Frequently asked questions

How much VRAM does Mistral Small 22B need?

Q4_K_M (a 13 GB download) needs 13 GB at a 4K context and 20 GB at its full 32,768-token context. That’s the model file, the memory its context needs and a 0.5 GB allowance for the program running it.

What’s the smallest GPU that can run Mistral Small 22B?

A graphics card with 16 GB of VRAM, such as the GeForce RTX 5080, GeForce RTX 5070 Ti and GeForce RTX 5060 Ti 16 GB, fits the Q4_K_M version at a 4K context with 2.5 GB to spare, and leaves room for up to 15,882 tokens of context. On a smaller card it can still run with part of the model in system RAM, but much more slowly.

Can Mistral Small 22B run on a Mac?

Yes, on a Mac with 24 GB of unified memory or more, such as the MacBook Air (M5), 24 GB, Mac mini (M6), 24 GB and Mac mini (M5 Pro), 24 GB. A Mac shares its memory with macOS and your apps, so we count about 18 GB of that 24 GB as usable.

How much memory does Mistral Small 22B’s context take?

About 219 MB for every 1,000 tokens, worked out from Mistral AI’s own config file. At its full 32,768-token context that adds 7 GB on top of the model file.

How fast will Mistral Small 22B run?

Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show the memory it needs and where it fits.

Did this tool do what you needed?

© 2026 Samprix. Not affiliated with OpenAI, Anthropic or Google.