Fine-tune your own language model
Fine-tuning teaches an open model your own vocabulary, tone or task: answering in Nepali, following your support scripts, writing in your house style. With LoRA or QLoRA you train a small set of extra weights instead of the whole model, which is why one rented GPU is enough for 7B–14B models.
The usual stack is PyTorch with Hugging Face Transformers, PEFT and TRL, or a ready-made trainer such as Axolotl or Unsloth. Start the PyTorch + JupyterLab software, upload your dataset to /workspace, and train.
Which GPU to pick
A 24 GB card such as the RTX 3090 or RTX 4090 handles QLoRA on 7B–14B models. 48 GB lets you train 30B-class models with QLoRA, or use full LoRA with longer context. 80 GB cards such as the A100 and H100 are fastest and suit 70B-class QLoRA.
Fine-tune a 7B model (LoRA, about 10,000 examples, 3 epochs) costs, on today’s cheapest hosts:
| GPU | Typical time | Cost today |
|---|---|---|
| RTX 3090 | 5 h | USD 1.15 |
| A100 SXM4 | 2 h | USD 1.56 |
| RTX 4090 | 3 h | USD 1.74 |
| A100 PCIE | 2.2 h | USD 1.80 |
| RTX A5000 | 5.5 h | USD 2.09 |
Ready-made software
Pick one of these when you rent and it starts with everything installed:
- PyTorch + JupyterLab — CUDA, PyTorch and JupyterLab ready — training, fine-tuning and notebooks. TensorBoard on port 6006.
- TensorFlow + JupyterLab — Official TensorFlow build with CUDA, JupyterLab and TensorBoard.
- CUDA dev box (bare) — Plain Ubuntu with CUDA + cuDNN, JupyterLab and SSH. Install whatever you like.
How it works
- Create an account with Google or your email address. There is no ID upload: your first payment is how we know who you are.
- Add credits in your own currency.
- Add your SSH key, then pick the software and a GPU from the live board.
- Wait for Running — about a minute on a warm host, a few minutes on a cold one — then open the app or connect over SSH.
- Stop when you are done. GPU billing ends at once; only the disk is billed until you destroy the instance.
Questions
How long does it take to fine-tune a 7B model?
A LoRA run on about 10,000 examples for three epochs typically takes two to five hours on a 24 GB card and about an hour on an H100. The table on this page prices exactly that job on today’s cheapest hosts.
Can I fine-tune a model to speak Nepali?
Yes. Llama, Qwen and Gemma have all seen some Nepali, and fine-tuning on your own Nepali examples is the usual way to make them better at it. Keep your dataset in Unicode Devanagari.
What happens to my model when I stop the instance?
Your files stay on the instance’s disk while it is stopped, and only storage is billed. Download the adapter or push it to Hugging Face before you destroy the instance, or turn on backups.