Samprix

How much VRAM does Mistral Small 24B need?

At a 4K context, Mistral Small 24B needs 14 GB of GPU memory with the Q4_K_M version. The smallest graphics card it fits on entirely has 16 GB of VRAM.

Mistral AI Specs last checked: how
Model
Quantization
Context length
Memory needed 14 GB Mistral Small 24B · Q4_K_M · 4,096 tokens
  • Model file 13 GB listed as 14 GB
  • Context 0.6 GB
  • Allowance 0.5 GB
File size from Ollama library: mistral-small; context memory from the publisher’s config, checked Oct 2, 2026.

Memory by context length

+156 MB per 1K tokens

Where it fits at 4K context

Memory available

Graphics cards

Macs

A Mac gives a model about 75% of its unified memory (our estimate). A card that’s short can still run it with part in system RAM, much more slowly. Check your machine

Frequently asked questions

How much VRAM does Mistral Small 24B need?

Q4_K_M (a 14 GB download) needs 14 GB at a 4K context and 19 GB at its full 32,768-token context. That’s the model file, the memory its context needs and a 0.5 GB allowance for the program running it.

What’s the smallest GPU that can run Mistral Small 24B?

A graphics card with 16 GB of VRAM, such as the GeForce RTX 5080, GeForce RTX 5070 Ti and GeForce RTX 5060 Ti 16 GB, fits the Q4_K_M version at a 4K context with 1.8 GB to spare, and leaves room for up to 16,131 tokens of context. On a smaller card it can still run with part of the model in system RAM, but much more slowly.

Can Mistral Small 24B run on a Mac?

Yes, on a Mac with 24 GB of unified memory or more, such as the MacBook Air (M5), 24 GB, Mac mini (M6), 24 GB and Mac mini (M5 Pro), 24 GB. A Mac shares its memory with macOS and your apps, so we count about 18 GB of that 24 GB as usable.

How much memory does Mistral Small 24B’s context take?

About 156 MB for every 1,000 tokens, worked out from Mistral AI’s own config file. At its full 32,768-token context that adds 5 GB on top of the model file.

How fast will Mistral Small 24B run?

Hardware makers and model publishers don’t publish tokens-per-second figures, so there is no source we could cite. We only show the memory it needs and where it fits.

Did this tool do what you needed?

© 2026 Samprix. Not affiliated with OpenAI, Anthropic or Google.