Which GPU do I need for fine-tuning?
The short answer: choose by GPU memory first. Speed matters, but a card that cannot hold your model cannot fine-tune it at any speed.
Why memory comes first
During fine-tuning the GPU has to hold the model's weights, the gradients, the optimiser's working data and your batch, all at once. That is more than the same model needs when it is only answering questions.
A rough way to estimate
Model weights take roughly 2GB per billion parameters at 16-bit precision, about 1GB at 8-bit and a little over 0.5GB at 4-bit. Fine-tuning adds to that. Parameter-efficient methods such as LoRA and QLoRA train a small set of extra weights on top of a frozen, often quantised, model, which is why they fit on far smaller cards than full fine-tuning does.
What that means for the cards we sell
| GPU memory | Card | Typical use |
|---|---|---|
| 16GB | RTX PRO 2000 | Learning, small models, QLoRA on smaller open models |
| 24GB | RTX PRO 4000 | QLoRA on mid-sized open models, development work |
| 32GB | RTX PRO 4500 | More headroom for batch size and context length |
| 48GB | RTX PRO 5000 | LoRA on larger models, small full fine-tunes |
| 72GB | RTX PRO 5000 72GB | Larger models again, or bigger batches |
| 96GB | RTX PRO 6000 | The largest models you can fine-tune on one workstation card |
These are starting points, not guarantees. Batch size, context length, precision and the software you use all change the answer.
Three questions to ask yourself
- What is the largest model you realistically plan to fine-tune in the next year?
- Will you use LoRA or QLoRA, or do you need full fine-tuning?
- Does your workstation have the power supply and space for the card?
If you can answer those, we can help you narrow it to one or two cards in a short call. See our Train and fine-tune solution for the hardware we recommend.