Can the GeForce RTX 3060 12GB (used) run Codestral 22B v0.1?
Only with CPU offload at Q4_K_M.
Codestral 22B v0.1 is 22.2B parameters. At Q4_K_M the weights are about 12.4 GiB; with the 1.5 GiB runtime reserve that is 13.9 GiB against the card's 12GB, about 1.9 GiB more than the card has.
All three quants
12GB card| Quant | Weights | With reserve | Verdict | Headroom |
|---|---|---|---|---|
| Q4_K_Mthe everyday quant | 12.4 GiB | 13.9 GiB | Only with CPU offload | short 1.9 GiB |
| Q8_0near-lossless | 21.9 GiB | 23.4 GiB | Only with CPU offload | short 11.4 GiB |
| FP16full weights | 41.4 GiB | 42.9 GiB | Does not fit | short 30.9 GiB |
Verdict rules: fits comfortably when weights plus reserve sit within 85% of VRAM; tight when they fit with little left for context; CPU offload when they exceed VRAM by up to 2x (runs, slowly, with layers in system RAM); no beyond that. The KV cache per token depends on the model's architecture and is not modelled here.
Not this card. These fit it comfortably at Q4.
- GeForce RTX 5090 32GB · headroom 18.1 GiB
- GeForce RTX 3090 (used) 24GB · headroom 10.1 GiB
- GeForce RTX 4090 (used) 24GB · headroom 10.1 GiB
- Radeon RX 7900 XTX 24GB · headroom 10.1 GiB
What the GeForce RTX 3060 12GB (used) does run well → · Codestral 22B v0.1 on every card →
About the model
Codestral 22B v0.1 by Mistral AI: 32,768-token native context, text, released 2024 under the Mistral AI Non-Production License. Model card →