picksbycard
August 23, 2026No. 113
Home / Will it run? / Qwen3 30B-A3B
Alibaba Qwen · Qwen3

Will Qwen3 30B-A3B run on my GPU?

30.5B parameters (mixture of experts, 3.3B active per token, but every expert must sit in memory), 32,768-token context, Apache-2.0. At Q4_K_M the weights come to about 17.0 GiB; add 1.5 GiB for the runtime and you need roughly 18.5 GiB of VRAM before context.

Weights by quant

GiB
Q4_K_M · the everyday quant17.0 GiB
Q8_0 · near-lossless30.1 GiB
FP16 · full weights56.8 GiB
Runtime reserve1.5 GiB

Card by card

4 of 17 fit at Q4
CardVRAMQ4_K_MQ8_0FP16Headroom at Q4
GeForce RTX 5090NVIDIA 32GB YesTightOffload 13.5 GiB
GeForce RTX 3090 (used)NVIDIA · used market 24GB YesOffloadNo 5.5 GiB
GeForce RTX 4090 (used)NVIDIA · used market 24GB YesOffloadNo 5.5 GiB
Radeon RX 7900 XTXAMD 24GB YesOffloadNo 5.5 GiB
GeForce RTX 5060 Ti 16GBNVIDIA 16GB OffloadOffloadNo short by 2.5 GiB
GeForce RTX 5070 TiNVIDIA 16GB OffloadOffloadNo short by 2.5 GiB
GeForce RTX 5080NVIDIA 16GB OffloadOffloadNo short by 2.5 GiB
Radeon RX 7800 XTAMD 16GB OffloadOffloadNo short by 2.5 GiB
Radeon RX 9060 XT 16GBAMD 16GB OffloadOffloadNo short by 2.5 GiB
Radeon RX 9070AMD 16GB OffloadOffloadNo short by 2.5 GiB
Radeon RX 9070 XTAMD 16GB OffloadOffloadNo short by 2.5 GiB
Arc B580Intel 12GB OffloadNoNo short by 6.5 GiB
GeForce RTX 3060 12GB (used)NVIDIA · used market 12GB OffloadNoNo short by 6.5 GiB
GeForce RTX 4070 SuperNVIDIA · used market 12GB OffloadNoNo short by 6.5 GiB
GeForce RTX 5070NVIDIA 12GB OffloadNoNo short by 6.5 GiB
Arc B570Intel 10GB OffloadNoNo short by 8.5 GiB
GeForce RTX 5060NVIDIA 8GB NoNoNo short by 10.5 GiB

Headroom is what remains for the KV cache after weights and the 1.5 GiB reserve. How far it stretches depends on the architecture: grouped-query models are frugal, older dense models are not. If the table says tight, plan on a shorter context or a lower quant. Parameter count from the model's own card.

The Silicon Note

One short read on the GPU market, most mornings. No spam, one-click out.