How much VRAM do you need for local AI?
Every number in this table comes from the people who publish the model or the tool, and each one is linked below. We only added one of our own: what SDXL actually used on our 12 GB card.
The rule of thumb
A model has to fit in the graphics card’s memory, plus room to work. For chat models that extra room holds the conversation (the longer the chat, the more it takes). For image models it holds the picture while it’s being made. So when a model’s file is 9.3 GB, a 12 GB card will run it, and a 16 GB card will run it with space for long chats.
If a model doesn’t fit, most tools don’t crash — they move part of it to normal system memory, and everything gets many times slower. That’s the thing you’re paying to avoid.
Chat models (LLMs)
Sizes are Ollama’s downloads of Qwen3, a popular open model family, in the default 4-bit version (q4_K_M) unless noted.
| Model | Download | 12 GB | 16 GB | 32 GB | 96 GB* |
|---|---|---|---|---|---|
| qwen3:4b | 2.5 GB | Yes | Yes | Yes | Yes |
| qwen3:8b | 5.2 GB | Yes | Yes | Yes | Yes |
| qwen3:8b (8-bit, q8_0) | 8.9 GB | Tight | Yes | Yes | Yes |
| qwen3:14b | 9.3 GB | Tight | Yes | Yes | Yes |
| qwen3:30b-a3b (mixture-of-experts) | 19 GB | No | No | Yes | Yes |
| qwen3:32b | 20 GB | No | No | Yes | Yes |
| qwen3:235b | 142 GB | No | No | No | No |
* 96 GB is what AMD lets a Ryzen AI Max+ 395 mini PC with 128 GB assign to graphics. “Tight” means it loads, with little room left for long conversations.
Speech to text
From OpenAI’s own Whisper table (their reference implementation; faster reimplementations like faster-whisper use less).
| Whisper model | Parameters | VRAM needed |
|---|---|---|
| small | 244 M | ~2 GB |
| medium | 769 M | ~5 GB |
| turbo | 809 M | ~6 GB |
| large | 1,550 M | ~10 GB |
We run a large-v3-turbo model for voice alignment on our 12 GB card. It’s fine alone — the trouble starts when it shares the card with an image model (see field notes).
Image generation
| Model | Main file | What it means |
|---|---|---|
| SDXL base 1.0 | 6.94 GB | Runs on 12 GB. On ours it peaked around 10 GB with the “VAE tiling” setting on; without it, the last step asked for about 17 GB and spilled into system memory. |
| FLUX.1 [dev], FP8 version | 12.33 GB | The main file alone is bigger than a 12 GB card. Aim for 16 GB or more. |
| FLUX.1 [dev], full (BF16) | 23.8 GB | A 32 GB card territory. Its text encoders come on top, although tools can move those off the card. |
So which size should you buy?
- 12 GB: 8B chat models, SDXL with care, Whisper. Workable — it’s what we use — but you’ll plan around it.
- 16 GB: 14B chat models with room, SDXL easily, FLUX in its smaller versions. The sensible starting point in 2026. See the buying guide.
- 32 GB: 30B–32B chat models, full-size FLUX.
- 96 GB (shared memory): models no single gaming card holds, at lower speed. See GPU vs mini PC.