Home / Will it run? / GLM-4.7-Flash
Z.ai (Zhipu) · GLM-4.7
Will GLM-4.7-Flash run on my GPU?
30B parameters (mixture of experts, 3B active per token, but every expert must sit in memory), 202,752-token context, MIT. At Q4_K_M the weights come to about 16.8 GiB; add 1.5 GiB for the runtime and you need roughly 18.3 GiB of VRAM before context.
Weights by quant
GiBQ4_K_M · the everyday quant16.8 GiB
Q8_0 · near-lossless29.6 GiB
FP16 · full weights55.9 GiB
Runtime reserve1.5 GiB
Card by card
4 of 17 fit at Q4| Card | VRAM | Q4_K_M | Q8_0 | FP16 | Headroom at Q4 |
|---|---|---|---|---|---|
| GeForce RTX 5090NVIDIA | 32GB | Yes | Tight | Offload | 13.7 GiB |
| GeForce RTX 3090 (used)NVIDIA · used market | 24GB | Yes | Offload | No | 5.7 GiB |
| GeForce RTX 4090 (used)NVIDIA · used market | 24GB | Yes | Offload | No | 5.7 GiB |
| Radeon RX 7900 XTXAMD | 24GB | Yes | Offload | No | 5.7 GiB |
| GeForce RTX 5060 Ti 16GBNVIDIA | 16GB | Offload | Offload | No | short by 2.3 GiB |
| GeForce RTX 5070 TiNVIDIA | 16GB | Offload | Offload | No | short by 2.3 GiB |
| GeForce RTX 5080NVIDIA | 16GB | Offload | Offload | No | short by 2.3 GiB |
| Radeon RX 7800 XTAMD | 16GB | Offload | Offload | No | short by 2.3 GiB |
| Radeon RX 9060 XT 16GBAMD | 16GB | Offload | Offload | No | short by 2.3 GiB |
| Radeon RX 9070AMD | 16GB | Offload | Offload | No | short by 2.3 GiB |
| Radeon RX 9070 XTAMD | 16GB | Offload | Offload | No | short by 2.3 GiB |
| Arc B580Intel | 12GB | Offload | No | No | short by 6.3 GiB |
| GeForce RTX 3060 12GB (used)NVIDIA · used market | 12GB | Offload | No | No | short by 6.3 GiB |
| GeForce RTX 4070 SuperNVIDIA · used market | 12GB | Offload | No | No | short by 6.3 GiB |
| GeForce RTX 5070NVIDIA | 12GB | Offload | No | No | short by 6.3 GiB |
| Arc B570Intel | 10GB | Offload | No | No | short by 8.3 GiB |
| GeForce RTX 5060NVIDIA | 8GB | No | No | No | short by 10.3 GiB |
Headroom is what remains for the KV cache after weights and the 1.5 GiB reserve. How far it stretches depends on the architecture: grouped-query models are frugal, older dense models are not. If the table says tight, plan on a shorter context or a lower quant. Parameter count from the model's own card.