Which models fit which card.
50 open-weight models against the 17 cards in our catalog. Weights are sized from the parameter count (all experts count for MoE), a 1.5 GiB reserve covers the runtime and a small KV cache, and what is left is your context headroom. Tap any cell for the full verdict at Q4, Q8 and FP16.
The grid
Q4_K_M verdictsMethod: weights in GiB = parameters x bytes per parameter (Q4_K_M 0.60, Q8_0 1.06, FP16 2.0), plus a fixed 1.5 GiB reserve. "Fits comfortably" means that total sits within 85% of the card's memory; "tight" means it fits with little room for context; "CPU offload" means the weights exceed VRAM by up to 2x and can still run slowly with layers offloaded to system RAM. Parameter counts come from each model's card or release notes; the per-token KV cache depends on the architecture and is not in this table, so treat the headroom figures as a budget, not a promise.