Courses / On-demand courses
Fine-Tuning LLMs
SFT, DPO and LoRA on open weights, with evals to prove it helped.
- recorded hours
- 9.5
- modules, in order
- 6
- deployed projects
- 4
- graded checkpoints
- 5
Who it is for, and what you receive.
Developers with Python and some PyTorch who want to fine-tune open weights and prove it helped. A GPU budget is stated up front.
Before you start
Python and some PyTorch; a GPU budget we specify.
Skills
- PyTorch
- LoRA
- DPO
- evals
- serving
- 1
The recorded sessions
Taught by an instructor who does the work, taken in order at your own pace, yours for twelve months.
- 2
Graded work after every module
Exercises and projects marked by Nova against a rubric you can read, with what you did well and what to fix.
- 3
The projects
Deployed and graded, each one on your record with its repository.
- 4
Your hiring plan
The same six-part plan the bootcamps use, run with the career agent.
- 5
The record and the certificate
Every grade at square1ai.com/u/{handle}; a credential ID that resolves at square1ai.com/verify.
- 6
Nova as your tutor
Help that is about your actual work, because it has read all of it.
The content plan, module by module.
9.5 recorded hours from an instructor, taken in order at your own pace. Nova grades the work at the end of each module.
- 1
When to fine-tune
- Prompting versus fine-tuning
- Task types
- Budgets
Graded
Nothing to submit; watch and take notes
60 min
- 2
Data for post-training
- Curation
- Synthetic data
- Provenance and licences
Graded
A curated training set with provenance
90 min
- 3
Supervised fine-tuning with LoRA
- LoRA
- Training loops
- Evals during training
Graded
A fine-tuned model beating the baseline
120 min
- 4
Preference tuning
- DPO
- Preference sets
- Win rate
Graded
A DPO run with a win-rate eval
90 min
- 5
Serving and cost
- Quantisation
- Serving
- Cost comparison
Graded
A cost comparison against a hosted API
90 min
- 6
Project: your model, deployed
- Deploy
- Monitor
- Report
Graded
The model in production with monitoring
120 min
The projects.
Each is deployed and graded by Nova against a rubric you can read before you start.
- 1
A training set
A curated set with provenance.
You hand in
- Set
- Provenance
The rubric requires
Every example has a source and a licence.
- 2
An SFT model
Beat the baseline.
You hand in
- Model
- Eval
The rubric requires
The gain is measured on a held-out set.
- 3
A preference-tuned model
Prefer the right answers.
You hand in
- DPO run
- Win-rate eval
The rubric requires
Win rate measured.
- 4
A deployed model
Serve it and cost it.
You hand in
- Serving
- Monitoring
- Cost comparison
The rubric requires
Deployed with a cost report.
What you can do at the end.
A post-training method from data to serving, and four graded projects on your record.
- 1
Decide when fine-tuning beats prompting.
- 2
Build a training set with provenance.
- 3
Fine-tune with LoRA and beat a baseline.
- 4
Preference-tune with DPO and measure win rate.
- 5
Serve the model and compare its cost to a hosted API.
Roles this prepares you for
- ML engineer (LLM)
- AI engineer
No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.
Your record at the end.
Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.
- Graded, line by line
- Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
- Module by module
- Each module ends in graded work; the next opens when you are ready, on your own schedule.
- Nova remembers
- Help is about your actual work, because the tutor has every submission and every failed exercise of yours.
- One page an employer can run
- Every grade and project at /verify. An employer opens it and runs the code.
Record, Fine-Tuning LLMs
Example
- ProjectGraded
A training set
Every example has a source and a licence.
- ProjectGraded
An SFT model
The gain is measured on a held-out set.
- ProjectGraded
A preference-tuned model
Win rate measured.
- ProjectGraded
A deployed model
Deployed with a cost report.
- Every moduleGraded
5 graded checkpoints
Each module ends in work Nova grades line by line against a rubric you can read.
- AfterKept
Your hiring plan
Target roles, the gap map, the proof to send, weekly actions and an interview log.
The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.
6 entries, at your own pace. One email when this course opens.
How we help you find a job.
Proof, not a certificate: the projects you deployed are the thing you show, and the tools below are yours to use.
- 1
Your hiring plan
The same six-part plan the bootcamps use: target roles, the gap map, the proof to send, weekly actions, an interview log, the outcome. You run it with the career agent.
- 2
A record an employer can run
Your graded projects on /verify. An employer opens it and runs the code.
- 3
The career agent
Paste a real job posting at /career and it maps the role to your graded work and what to do next.
- 4
The roles directory
Every role we prepare people for, what it pays and what it asks, at /roles.
- 5
A path to the live cohort
If you want the instructor, the gates and the hiring sprint, the bootcamp on the same subject is one waitlist away.
One email when this course opens.
Recorded by an instructor who does this work, graded by Nova. The waitlist hears the date and the price first, and nothing is charged before it exists.
