Courses / On-demand courses
Evaluating AI Systems
LLM judges, rubrics and regression suites you can trust.
- recorded hours
- 7
- modules, in order
- 5
- deployed projects
- 3
- graded checkpoints
- 5
Who it is for, and what you receive.
Developers in Python who have shipped or tested an LLM feature and need to say how good it is, with evidence.
Before you start
Python; you have shipped or tested an LLM feature.
Skills
- evals
- LLM judges
- CI
- labelling
- 1
The recorded sessions
Taught by an instructor who does the work, taken in order at your own pace, yours for twelve months.
- 2
Graded work after every module
Exercises and projects marked by Nova against a rubric you can read, with what you did well and what to fix.
- 3
The projects
Deployed and graded, each one on your record with its repository.
- 4
Your hiring plan
The same six-part plan the bootcamps use, run with the career agent.
- 5
The record and the certificate
Every grade at square1ai.com/u/{handle}; a credential ID that resolves at square1ai.com/verify.
- 6
Nova as your tutor
Help that is about your actual work, because it has read all of it.
The content plan, module by module.
7 recorded hours from an instructor, taken in order at your own pace. Nova grades the work at the end of each module.
- 1
Defining quality
- Rubrics
- Agreement
- What to measure
Graded
A rubric two people agree on
60 min
- 2
Golden sets
- Sampling
- Labelling
- Maintaining the set
Graded
A labelled golden set
90 min
- 3
LLM judges, calibrated
- Judge prompts
- Calibration
- Drift
Graded
A judge with measured agreement
90 min
- 4
Regression suites in CI
- Suites
- Thresholds
- Blocking merges
Graded
An eval pipeline that blocks a bad change
90 min
- 5
Project: evals for a real feature
- Set
- Judge
- Pipeline
Graded
The suite, running
90 min
The projects.
Each is deployed and graded by Nova against a rubric you can read before you start.
- 1
A golden set and rubric
A rubric and 200 labelled cases.
You hand in
- Rubric
- Golden set
- Agreement report
The rubric requires
Agreement above the threshold.
- 2
A calibrated judge
A judge that agrees with the labellers.
You hand in
- Judge
- Calibration report
The rubric requires
Agreement measured.
- 3
A CI eval pipeline
Block a bad change.
You hand in
- Pipeline
- Thresholds
The rubric requires
A seeded regression is caught.
What you can do at the end.
You can say how good an AI feature is, with evidence, and three graded projects on your record.
- 1
Define quality as a rubric two people agree on.
- 2
Build a labelled golden set.
- 3
Calibrate an LLM judge against human labels.
- 4
Run regression suites in CI that block a bad change.
- 5
Build evals for a real feature.
Roles this prepares you for
- AI engineer
- QA engineer (AI)
No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.
Your record at the end.
Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.
- Graded, line by line
- Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
- Module by module
- Each module ends in graded work; the next opens when you are ready, on your own schedule.
- Nova remembers
- Help is about your actual work, because the tutor has every submission and every failed exercise of yours.
- One page an employer can run
- Every grade and project at /verify. An employer opens it and runs the code.
Record, Evaluating AI Systems
Example
- ProjectGraded
A golden set and rubric
Agreement above the threshold.
- ProjectGraded
A calibrated judge
Agreement measured.
- ProjectGraded
A CI eval pipeline
A seeded regression is caught.
- Every moduleGraded
5 graded checkpoints
Each module ends in work Nova grades line by line against a rubric you can read.
- AfterKept
Your hiring plan
Target roles, the gap map, the proof to send, weekly actions and an interview log.
The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.
5 entries, at your own pace. One email when this course opens.
How we help you find a job.
Proof, not a certificate: the projects you deployed are the thing you show, and the tools below are yours to use.
- 1
Your hiring plan
The same six-part plan the bootcamps use: target roles, the gap map, the proof to send, weekly actions, an interview log, the outcome. You run it with the career agent.
- 2
A record an employer can run
Your graded projects on /verify. An employer opens it and runs the code.
- 3
The career agent
Paste a real job posting at /career and it maps the role to your graded work and what to do next.
- 4
The roles directory
Every role we prepare people for, what it pays and what it asks, at /roles.
- 5
A path to the live cohort
If you want the instructor, the gates and the hiring sprint, the bootcamp on the same subject is one waitlist away.
One email when this course opens.
Recorded by an instructor who does this work, graded by Nova. The waitlist hears the date and the price first, and nothing is charged before it exists.
