Home / Will it run? / Gemma 4 26B-A4B IT
Google · Gemma 4
Will Gemma 4 26B-A4B IT run on my GPU?
25.8B parameters (mixture of experts, 3.8B active per token, but every expert must sit in memory), 262,144-token context, Apache-2.0. At Q4_K_M the weights come to about 14.4 GiB; add 1.5 GiB for the runtime and you need roughly 15.9 GiB of VRAM before context.
Weights by quant
GiBQ4_K_M · the everyday quant14.4 GiB
Q8_0 · near-lossless25.5 GiB
FP16 · full weights48.1 GiB
Runtime reserve1.5 GiB
Card by card
11 of 17 fit at Q4Headroom is what remains for the KV cache after weights and the 1.5 GiB reserve. How far it stretches depends on the architecture: grouped-query models are frugal, older dense models are not. If the table says tight, plan on a shorter context or a lower quant. Parameter count from the model's own card.