Which GPU do I need for fine-tuning?

The short answer: choose by GPU memory first. Speed matters, but a card that cannot hold your model cannot fine-tune it at any speed.

Why memory comes first

During fine-tuning the GPU has to hold the model's weights, the gradients, the optimiser's working data and your batch, all at once. That is more than the same model needs when it is only answering questions.

A rough way to estimate

Model weights take roughly 2GB per billion parameters at 16-bit precision, about 1GB at 8-bit and a little over 0.5GB at 4-bit. Fine-tuning adds to that. Parameter-efficient methods such as LoRA and QLoRA train a small set of extra weights on top of a frozen, often quantised, model, which is why they fit on far smaller cards than full fine-tuning does.

What that means for the cards we sell

GPU memory Card Typical use
16GB RTX PRO 2000 Learning, small models, QLoRA on smaller open models
24GB RTX PRO 4000 QLoRA on mid-sized open models, development work
32GB RTX PRO 4500 More headroom for batch size and context length
48GB RTX PRO 5000 LoRA on larger models, small full fine-tunes
72GB RTX PRO 5000 72GB Larger models again, or bigger batches
96GB RTX PRO 6000 The largest models you can fine-tune on one workstation card

These are starting points, not guarantees. Batch size, context length, precision and the software you use all change the answer.

Three questions to ask yourself

  • What is the largest model you realistically plan to fine-tune in the next year?
  • Will you use LoRA or QLoRA, or do you need full fine-tuning?
  • Does your workstation have the power supply and space for the card?

If you can answer those, we can help you narrow it to one or two cards in a short call. See our Train and fine-tune solution for the hardware we recommend.