TechWise Co

How much VRAM do you need for local AI?

By Roger Chacón · Updated

Every number in this table comes from the people who publish the model or the tool, and each one is linked below. We only added one of our own: what SDXL actually used on our 12 GB card.

The rule of thumb

A model has to fit in the graphics card’s memory, plus room to work. For chat models that extra room holds the conversation (the longer the chat, the more it takes). For image models it holds the picture while it’s being made. So when a model’s file is 9.3 GB, a 12 GB card will run it, and a 16 GB card will run it with space for long chats.

If a model doesn’t fit, most tools don’t crash — they move part of it to normal system memory, and everything gets many times slower. That’s the thing you’re paying to avoid.

Chat models (LLMs)

Sizes are Ollama’s downloads of Qwen3, a popular open model family, in the default 4-bit version (q4_K_M) unless noted.

ModelDownload12 GB16 GB32 GB96 GB*
qwen3:4b2.5 GBYesYesYesYes
qwen3:8b5.2 GBYesYesYesYes
qwen3:8b (8-bit, q8_0)8.9 GBTightYesYesYes
qwen3:14b9.3 GBTightYesYesYes
qwen3:30b-a3b (mixture-of-experts)19 GBNoNoYesYes
qwen3:32b20 GBNoNoYesYes
qwen3:235b142 GBNoNoNoNo

* 96 GB is what AMD lets a Ryzen AI Max+ 395 mini PC with 128 GB assign to graphics. “Tight” means it loads, with little room left for long conversations.

Speech to text

From OpenAI’s own Whisper table (their reference implementation; faster reimplementations like faster-whisper use less).

Whisper modelParametersVRAM needed
small244 M~2 GB
medium769 M~5 GB
turbo809 M~6 GB
large1,550 M~10 GB

We run a large-v3-turbo model for voice alignment on our 12 GB card. It’s fine alone — the trouble starts when it shares the card with an image model (see field notes).

Image generation

ModelMain fileWhat it means
SDXL base 1.06.94 GBRuns on 12 GB. On ours it peaked around 10 GB with the “VAE tiling” setting on; without it, the last step asked for about 17 GB and spilled into system memory.
FLUX.1 [dev], FP8 version12.33 GBThe main file alone is bigger than a 12 GB card. Aim for 16 GB or more.
FLUX.1 [dev], full (BF16)23.8 GBA 32 GB card territory. Its text encoders come on top, although tools can move those off the card.

So which size should you buy?

Sources

  1. Ollama — qwen3 tags and download sizes
  2. OpenAI Whisper — available models and VRAM
  3. Stability AI — SDXL base 1.0 files
  4. Black Forest Labs — FLUX.1 [dev] and FLUX.1 [dev] FP8
  5. AMD — Variable Graphics Memory FAQ