Skip to content
LLMs & Agents

When should you fine-tune an open-weights model instead of using an API?

Fine-tune an open-weights model only after prompting and retrieval are measured and still short, and when the gap is format consistency, a narrow high-volume task, latency, or data that must stay in-house. Fine-tuning appeared in 24 of 315 AI job ads (8%) in our 19 September 2026 sample, against evals in 23%.

Nikhil De Silva · Founder, Square 1 AI6 min read

Fine-tune an open-weights model when you have already tried prompting and retrieval, measured where they fall short, and the gap is one fine-tuning is good at closing: a consistent format or style, a narrow task done at high volume, a cost or latency target a hosted API cannot meet, or data that must not leave your own infrastructure. If you cannot say which of those applies, and show it with an eval, use the API. Fine-tuning is a real skill with a real hosting bill attached, not an upgrade you apply because the model "should know our domain".

Employers treat it the same way. In our sample of 315 AI and machine learning job ads collected on 19 September 2026, 24 (8%) mentioned fine-tuning. Evals appeared in 73 (23%) and retrieval or RAG in 34 (11%). Fine-tuning is a specialism a minority of teams hire for; knowing when not to do it is the more common requirement.

What should you try before fine-tuning?

In this order, and measure each step on the same test set:

  1. A better prompt. Clear instructions, a defined output format, a few worked examples. Many "we need to fine-tune" projects end here. Our comparison of fine-tuning and prompt engineering goes through where each wins.
  2. Retrieval. If the model is wrong because it lacks facts, your documents, your policies, last week's prices, fine-tuning is the wrong tool: it teaches behaviour, not up-to-date knowledge. Retrieval puts the facts in front of the model at question time and lets you cite them. See RAG versus fine-tuning.
  3. Structured output and tools. Schema-constrained output and a validation step fix many "the format keeps breaking" complaints without training anything.
  4. A bigger or different hosted model. Sometimes the cheapest fix is a stronger model for the hard cases and a small one for the easy ones.

Only when all of that is measured and still short does fine-tuning earn a place.

When does fine-tuning actually pay?

Five situations, and most real projects that succeed have two or more of them:

Situation Why fine-tuning helps What to check first
Format or style must be consistent The behaviour moves from a long prompt into the weights, so it holds without reminders Whether structured output already solves it
A narrow task at high volume A small tuned model can match a large general one on one task, at lower cost per call Your real call volume, not a hoped-for one
Latency matters A smaller model on your own hardware answers faster and more predictably Where the time actually goes today
Data must stay in-house An open-weights model can run in your cloud or on your premises Whether a hosted provider's data terms would satisfy your security team
The hosted models do the task badly Specialised labels, jargon or a rare language the general models fumble Whether the problem is really missing knowledge, which is retrieval's job

The common thread is a narrow, stable task. Fine-tuning rewards a job that will look the same next month; it punishes one whose requirements keep moving, because every change means new data and a new training run.

What are LoRA and QLoRA, in plain terms?

Full fine-tuning updates every weight in the model, which needs a lot of GPU memory and produces a full copy of the model for every version. LoRA (low-rank adaptation) freezes the original weights and trains a small set of extra weights alongside them. The result is a small adapter file you load on top of the base model, so training is cheaper, you can keep several adapters for different tasks, and rolling back is deleting a file. QLoRA does the same on a base model held in a compressed, quantised form, which cuts memory further and lets people tune larger models on a single GPU, at some cost in training speed.

For most teams starting out, a LoRA fine-tune of a small or mid-sized open model on a modest set of well-chosen examples is the sensible first experiment. Preference tuning (DPO and its relatives), where you show the model pairs of better and worse answers, is the usual next step once supervised fine-tuning has plateaued.

How do you know the fine-tune worked?

Build the eval before you train anything. The sequence that keeps you honest:

  • A held-out test set drawn from real inputs, never shown to the model during training.
  • A baseline: the best prompted and retrieval-backed result on that set, with a hosted model and with the untuned open model.
  • The same metric after training, on the same set. If the gain is not there, you have your answer.
  • Regression checks on things the model used to do well, because tuning on a narrow task can make it worse at others.
  • Evals during training, so you can see overfitting when the training loss keeps falling and the held-out score does not.

The data matters more than the method. A smaller set of clean, representative, correctly licensed examples beats a large scraped one, and you should know where every example came from.

What does it cost to host your own model?

This is the part people underestimate. A hosted API is someone else's GPUs, uptime, scaling and security patches. When you fine-tune an open model, those become yours: serving (often with an inference server and quantisation), autoscaling for peaks, monitoring for quality drift, upgrades when a better base model appears and you have to retrain, and an on-call rota. Compare cost per token honestly, including idle GPUs and engineer time, not just the rental price at full load. Our piece on managed versus self-hosted LLMs sets out the trade-off. A middle path: some providers fine-tune and host a model for you, trading some control for less operational load.

Where should you start this week?

Pick one narrow task you already run through an API. Collect a couple of hundred real inputs with correct outputs, hold a slice back as a test set, and score your current prompt on it. Then try the cheap fixes above and score again. If a gap remains and it is the kind in the table, run one small LoRA fine-tune of an open model on the rest and compare on the held-out slice. Either way, you keep the eval.

Where does Square 1 teach this?

The Fine-Tuning and Open-Weights Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, in six blocks that each end in a deployed project and a gate: a reproducible baseline of three open models on your task, a training set where every example has a source and a licence, a LoRA fine-tune that beats the baseline on a held-out set, a preference-tuned model with a measured win rate, a production deployment with monitoring and a cost comparison against a hosted API defended in a recorded viva, and an employer brief with a hiring sprint. It asks for Python and some PyTorch, and states a GPU budget up front: about A$150 over the twelve weeks. Fine-Tuning LLMs is the on-demand version of the method, recorded by an instructor and graded by Nova, the AI tutor, starting with when to fine-tune at all. Retrieval-Augmented Generation covers the retrieval route you should try first. All three are taking a waitlist. The machine learning engineer role page lists the day-to-day, and the free generative AI skill check takes about three minutes.

Questions people ask

Should I fine-tune or use RAG?

If the model is wrong because it lacks facts, use retrieval: fine-tuning teaches behaviour, not up-to-date knowledge. Fine-tuning suits consistent formats, narrow high-volume tasks, latency targets and data that must stay on your own infrastructure.

When is fine-tuning an open-weights model worth it?

When a better prompt, retrieval and structured output have been measured on a held-out test set and still fall short, and the task is narrow and stable. It pays most at high volume, under latency or privacy constraints, or where hosted models handle the task badly.

What is the difference between LoRA and QLoRA?

LoRA freezes the base model and trains a small adapter alongside it, so training is cheaper and versions are small files. QLoRA does the same on a quantised, compressed base model, cutting memory further so larger models can be tuned on a single GPU.

How often do AI job ads ask for fine-tuning?

In 315 AI and machine learning job ads collected on 19 September 2026, 24 (8%) mentioned fine-tuning, while 73 (23%) mentioned evals and 34 (11%) retrieval or RAG.

What does self-hosting a fine-tuned model involve?

Serving, autoscaling, quality monitoring, retraining when a better base model arrives, and on-call support all become your team's job. Compare cost per token including idle GPUs and engineer time, not only the rental price at full load.

Free skill check · about 3 minutes

Where do you stand on Generative AI?

Five questions, and a skill breakdown the moment you finish: your strengths, the gaps to close, and what to learn next from real curriculum.

Start the Generative AI skill check

Free, with a student account — the check is the first entry in your record.

Learn this by building it

The programmes that teach what this piece covers, each ending in deployed work graded against a rubric you can read.