Fine-tune an open-weights model when you have already tried prompting and retrieval, measured where they fall short, and the gap is one fine-tuning is good at closing: a consistent format or style, a narrow task done at high volume, a cost or latency target a hosted API cannot meet, or data that must not leave your own infrastructure. If you cannot say which of those applies, and show it with an eval, use the API. Fine-tuning is a real skill with a real hosting bill attached, not an upgrade you apply because the model "should know our domain".
Employers treat it the same way. In our sample of 315 AI and machine learning job ads collected on 19 September 2026, 24 (8%) mentioned fine-tuning. Evals appeared in 73 (23%) and retrieval or RAG in 34 (11%). Fine-tuning is a specialism a minority of teams hire for; knowing when not to do it is the more common requirement.
What should you try before fine-tuning?
In this order, and measure each step on the same test set:
- A better prompt. Clear instructions, a defined output format, a few worked examples. Many "we need to fine-tune" projects end here. Our comparison of fine-tuning and prompt engineering goes through where each wins.
- Retrieval. If the model is wrong because it lacks facts, your documents, your policies, last week's prices, fine-tuning is the wrong tool: it teaches behaviour, not up-to-date knowledge. Retrieval puts the facts in front of the model at question time and lets you cite them. See RAG versus fine-tuning.
- Structured output and tools. Schema-constrained output and a validation step fix many "the format keeps breaking" complaints without training anything.
- A bigger or different hosted model. Sometimes the cheapest fix is a stronger model for the hard cases and a small one for the easy ones.
Only when all of that is measured and still short does fine-tuning earn a place.
When does fine-tuning actually pay?
Five situations, and most real projects that succeed have two or more of them:
| Situation | Why fine-tuning helps | What to check first |
|---|---|---|
| Format or style must be consistent | The behaviour moves from a long prompt into the weights, so it holds without reminders | Whether structured output already solves it |
| A narrow task at high volume | A small tuned model can match a large general one on one task, at lower cost per call | Your real call volume, not a hoped-for one |
| Latency matters | A smaller model on your own hardware answers faster and more predictably | Where the time actually goes today |
| Data must stay in-house | An open-weights model can run in your cloud or on your premises | Whether a hosted provider's data terms would satisfy your security team |
| The hosted models do the task badly | Specialised labels, jargon or a rare language the general models fumble | Whether the problem is really missing knowledge, which is retrieval's job |
The common thread is a narrow, stable task. Fine-tuning rewards a job that will look the same next month; it punishes one whose requirements keep moving, because every change means new data and a new training run.
What are LoRA and QLoRA, in plain terms?
Full fine-tuning updates every weight in the model, which needs a lot of GPU memory and produces a full copy of the model for every version. LoRA (low-rank adaptation) freezes the original weights and trains a small set of extra weights alongside them. The result is a small adapter file you load on top of the base model, so training is cheaper, you can keep several adapters for different tasks, and rolling back is deleting a file. QLoRA does the same on a base model held in a compressed, quantised form, which cuts memory further and lets people tune larger models on a single GPU, at some cost in training speed.
For most teams starting out, a LoRA fine-tune of a small or mid-sized open model on a modest set of well-chosen examples is the sensible first experiment. Preference tuning (DPO and its relatives), where you show the model pairs of better and worse answers, is the usual next step once supervised fine-tuning has plateaued.
How do you know the fine-tune worked?
Build the eval before you train anything. The sequence that keeps you honest:
- A held-out test set drawn from real inputs, never shown to the model during training.
- A baseline: the best prompted and retrieval-backed result on that set, with a hosted model and with the untuned open model.
- The same metric after training, on the same set. If the gain is not there, you have your answer.
- Regression checks on things the model used to do well, because tuning on a narrow task can make it worse at others.
- Evals during training, so you can see overfitting when the training loss keeps falling and the held-out score does not.
The data matters more than the method. A smaller set of clean, representative, correctly licensed examples beats a large scraped one, and you should know where every example came from.
What does it cost to host your own model?
This is the part people underestimate. A hosted API is someone else's GPUs, uptime, scaling and security patches. When you fine-tune an open model, those become yours: serving (often with an inference server and quantisation), autoscaling for peaks, monitoring for quality drift, upgrades when a better base model appears and you have to retrain, and an on-call rota. Compare cost per token honestly, including idle GPUs and engineer time, not just the rental price at full load. Our piece on managed versus self-hosted LLMs sets out the trade-off. A middle path: some providers fine-tune and host a model for you, trading some control for less operational load.
Where should you start this week?
Pick one narrow task you already run through an API. Collect a couple of hundred real inputs with correct outputs, hold a slice back as a test set, and score your current prompt on it. Then try the cheap fixes above and score again. If a gap remains and it is the kind in the table, run one small LoRA fine-tune of an open model on the rest and compare on the held-out slice. Either way, you keep the eval.
Where does Square 1 teach this?
The Fine-Tuning and Open-Weights Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, in six blocks that each end in a deployed project and a gate: a reproducible baseline of three open models on your task, a training set where every example has a source and a licence, a LoRA fine-tune that beats the baseline on a held-out set, a preference-tuned model with a measured win rate, a production deployment with monitoring and a cost comparison against a hosted API defended in a recorded viva, and an employer brief with a hiring sprint. It asks for Python and some PyTorch, and states a GPU budget up front: about A$150 over the twelve weeks. Fine-Tuning LLMs is the on-demand version of the method, recorded by an instructor and graded by Nova, the AI tutor, starting with when to fine-tune at all. Retrieval-Augmented Generation covers the retrieval route you should try first. All three are taking a waitlist. The machine learning engineer role page lists the day-to-day, and the free generative AI skill check takes about three minutes.
