Skip to content

Courses / On-demand courses

On demandRecorded by an instructor, graded by Nova

Evaluating AI Systems

LLM judges, rubrics and regression suites you can trust.

recorded hours
7
modules, in order
5
deployed projects
3
graded checkpoints
5

Who it is for, and what you receive.

Developers in Python who have shipped or tested an LLM feature and need to say how good it is, with evidence.

Before you start

Python; you have shipped or tested an LLM feature.

Skills

  • evals
  • LLM judges
  • CI
  • labelling
  1. 1

    The recorded sessions

    Taught by an instructor who does the work, taken in order at your own pace, yours for twelve months.

  2. 2

    Graded work after every module

    Exercises and projects marked by Nova against a rubric you can read, with what you did well and what to fix.

  3. 3

    The projects

    Deployed and graded, each one on your record with its repository.

  4. 4

    Your hiring plan

    The same six-part plan the bootcamps use, run with the career agent.

  5. 5

    The record and the certificate

    Every grade at square1ai.com/u/{handle}; a credential ID that resolves at square1ai.com/verify.

  6. 6

    Nova as your tutor

    Help that is about your actual work, because it has read all of it.

The content plan, module by module.

7 recorded hours from an instructor, taken in order at your own pace. Nova grades the work at the end of each module.

  1. 1

    Defining quality

    • Rubrics
    • Agreement
    • What to measure

    Graded

    A rubric two people agree on

    60 min

  2. 2

    Golden sets

    • Sampling
    • Labelling
    • Maintaining the set

    Graded

    A labelled golden set

    90 min

  3. 3

    LLM judges, calibrated

    • Judge prompts
    • Calibration
    • Drift

    Graded

    A judge with measured agreement

    90 min

  4. 4

    Regression suites in CI

    • Suites
    • Thresholds
    • Blocking merges

    Graded

    An eval pipeline that blocks a bad change

    90 min

  5. 5

    Project: evals for a real feature

    • Set
    • Judge
    • Pipeline

    Graded

    The suite, running

    90 min

The projects.

Each is deployed and graded by Nova against a rubric you can read before you start.

  1. 1

    A golden set and rubric

    A rubric and 200 labelled cases.

    You hand in

    • Rubric
    • Golden set
    • Agreement report

    The rubric requires

    Agreement above the threshold.

  2. 2

    A calibrated judge

    A judge that agrees with the labellers.

    You hand in

    • Judge
    • Calibration report

    The rubric requires

    Agreement measured.

  3. 3

    A CI eval pipeline

    Block a bad change.

    You hand in

    • Pipeline
    • Thresholds

    The rubric requires

    A seeded regression is caught.

What you can do at the end.

You can say how good an AI feature is, with evidence, and three graded projects on your record.

  1. 1

    Define quality as a rubric two people agree on.

  2. 2

    Build a labelled golden set.

  3. 3

    Calibrate an LLM judge against human labels.

  4. 4

    Run regression suites in CI that block a bad change.

  5. 5

    Build evals for a real feature.

Roles this prepares you for

  • AI engineer
  • QA engineer (AI)

No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.

Your record at the end.

Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.

Graded, line by line
Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
Module by module
Each module ends in graded work; the next opens when you are ready, on your own schedule.
Nova remembers
Help is about your actual work, because the tutor has every submission and every failed exercise of yours.
One page an employer can run
Every grade and project at /verify. An employer opens it and runs the code.

Record, Evaluating AI Systems

Example

  1. Project

    A golden set and rubric

    Agreement above the threshold.

    Graded
  2. Project

    A calibrated judge

    Agreement measured.

    Graded
  3. Project

    A CI eval pipeline

    A seeded regression is caught.

    Graded
  4. Every module

    5 graded checkpoints

    Each module ends in work Nova grades line by line against a rubric you can read.

    Graded
  5. After

    Your hiring plan

    Target roles, the gap map, the proof to send, weekly actions and an interview log.

    Kept

The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.

5 entries, at your own pace. One email when this course opens.

No account, no card. One email when it opens; we never sell before it exists.

How we help you find a job.

Proof, not a certificate: the projects you deployed are the thing you show, and the tools below are yours to use.

  1. 1

    Your hiring plan

    The same six-part plan the bootcamps use: target roles, the gap map, the proof to send, weekly actions, an interview log, the outcome. You run it with the career agent.

  2. 2

    A record an employer can run

    Your graded projects on /verify. An employer opens it and runs the code.

  3. 3

    The career agent

    Paste a real job posting at /career and it maps the role to your graded work and what to do next.

  4. 4

    The roles directory

    Every role we prepare people for, what it pays and what it asks, at /roles.

  5. 5

    A path to the live cohort

    If you want the instructor, the gates and the hiring sprint, the bootcamp on the same subject is one waitlist away.

One email when this course opens.

Recorded by an instructor who does this work, graded by Nova. The waitlist hears the date and the price first, and nothing is charged before it exists.

No account, no card. One email when it opens; we never sell before it exists.