Skip to content

Courses / Live bootcamps

Live bootcampTaught live on Zoom, one instructor

LLM Evals and Reliability Bootcamp

Make an AI system you can trust: evals, guardrails, monitoring, incident review.

weeks, live on Zoom
12
hours a week
15
deployed projects
6
minutes 1-1, every week
30

Who it is for, and what you receive.

Engineers who have shipped an LLM feature and been burned by it, and teams that need someone to own quality.

Before you start

You have shipped an LLM feature and been burned by it.

Skills

  • Python
  • evals
  • LLM judges
  • CI
  • guardrails
  • monitoring
  • incident review

A week, about 15 hours

Live on Zoom
4h
Recorded lessons and exercises
6h
The project
5h
  1. 1

    Six deployed projects

    Each graded by Nova against a published rubric and signed off by the instructor at the gate.

  2. 2

    The recorded lessons

    The track's lessons and exercises, graded line by line, open for twelve months after the cohort ends.

  3. 3

    Live on Zoom every week

    A 90-minute live code review, a 60-minute squad lab, 60 minutes of office hours, and a 30-minute 1-1 with your instructor.

  4. 4

    Your hiring plan

    Opened in week 1 and reviewed in every 1-1: target roles, the gap map, the proof list, weekly actions, an interview log, the outcome.

  5. 5

    The record

    Every grade and every project with its repository at square1ai.com/u/{handle}, and the recording of your viva, verifiable by any employer.

  6. 6

    The certificate

    A credential ID that resolves at square1ai.com/verify to the real completion.

  7. 7

    The hiring sprint

    Weeks 11 and 12: CV and portfolio from graded work, applications, mock interviews scored against the role, demo day.

  8. 8

    Nova for twelve weeks

    A tutor with every submission of yours in its memory, at 2am as well as in class.

Twelve weeks in six blocks.

Each block teaches for a week, then you build and deploy a project, then a gate checks it before the next block opens. The project and its gate are the block's record entry.

Live time each week on Zoom: a 90-minute class where the instructor reviews real submissions, a 60-minute squad lab, 60 minutes of office hours, and your own 30-minute 1-1. Nothing is lectured live; the recorded lessons do that.

  1. 1

    What good looks like

    Weeks 1 to 2

    Week 1

    What good looks like

    • Defining quality
    • Rubrics and inter-rater agreement
    • Sampling and labelling
    • Golden sets

    You build: A rubric and a labelling plan

    Week 2

    Project 1: the golden set

    • Labelling 300 cases
    • Agreement reports

    You build: Labelled golden set, published

    Project 1, the gate at week 2

    Golden set

    Define quality for a support-answer feature and label 300 cases.

    You hand in

    • Rubric
    • Labelling guide
    • Golden set
    • Agreement report

    The gate

    Two labellers agree above the stated threshold.

  2. 2

    Judges

    Weeks 3 to 4

    Week 3

    Judges

    • LLM-as-judge
    • Pairwise and absolute scoring
    • Calibration
    • Judge drift

    You build: A calibrated judge

    Week 4

    Project 2: the judge

    • Judge as a service
    • Measuring agreement

    You build: Judge service with measured agreement

    Project 2, the gate at week 4

    Calibrated judge

    A judge that scores answers the way the labellers do.

    You hand in

    • Judge service
    • Calibration report
    • Agreement with humans

    The gate

    Judge agreement measured and reported.

  3. 3

    Regression and CI

    Weeks 5 to 6

    Week 5

    Regression and CI

    • Eval suites in the build
    • Prompt and model versioning
    • Flakiness
    • Sampling budgets

    You build: Eval pipeline; squads form

    Week 6

    Project 3: the pipeline

    • Seeded regressions
    • Change control

    You build: CI pipeline blocking a bad prompt change

    Project 3, the gate at week 6

    Regression in CI

    Make prompt and model changes safe.

    You hand in

    • Eval pipeline
    • Version control for prompts
    • Sampling policy

    The gate

    A seeded regression is caught before merge.

  4. 4

    Guardrails

    Weeks 7 to 8

    Week 7

    Guardrails

    • Input and output checks
    • Prompt injection defences
    • PII and policy
    • Evaluating the guardrails

    You build: Guardrail layer

    Week 8

    Project 4: the guardrails

    • Attack suites
    • Protecting the golden set

    You build: Guardrails deployed with their own evals

    Project 4, the gate at week 8

    Guardrails

    Stop the attacks without hurting the good cases.

    You hand in

    • Input and output guards
    • Injection suite
    • PII policy
    • Guardrail evals

    The gate

    The attack suite is blocked; the golden set is unharmed.

  5. 5

    Production monitoring

    Weeks 9 to 10

    Week 9

    Production monitoring

    • Online evals and sampling
    • Dashboards that mean something
    • Alerting
    • Incident review for AI

    You build: Monitoring design

    Week 10

    Project 5 and the viva

    • The monitored feature
    • Walking through an incident

    You build: Monitored feature with alerts and a written incident review; the viva

    Project 5, the gate at week 10

    Monitoring and incident

    Run the feature in production and survive an incident.

    You hand in

    • Online evals
    • Dashboard
    • Alerts
    • A written incident review

    The gate

    The viva: defend the monitoring and walk through the incident.

  6. 6

    Employer brief and hiring sprint

    Weeks 11 to 12

    Week 11

    Employer brief

    • A real problem from a hiring partner, in squads
    • A partner's reliability problem
    • Working to someone else's definition of done

    You build: The employer brief in progress

    Week 12

    Hiring sprint

    • CV and portfolio built from graded work
    • Applications and follow-ups in the hiring plan
    • Mock interviews scored against the role
    • Demo day

    You build: Project 6 delivered; demo day

    Project 6, the gate at week 12

    Employer brief

    A partner's reliability problem.

    You hand in

    • The brief delivered
    • Demo day

    The gate

    The brief's owner accepts the result.

What you can do by week 12.

The quality infrastructure every AI team needs and few have: a golden set, a calibrated judge, a CI pipeline, guardrails and monitoring, built once and carried to any team.

  1. 1

    Define quality for an LLM feature as a rubric people agree on, and build a labelled golden set.

  2. 2

    Build LLM judges calibrated against human labels and measure their agreement.

  3. 3

    Run eval suites in CI with change control for prompts and models, and handle flakiness and sampling.

  4. 4

    Build guardrails for inputs, outputs, prompt injection and PII, and evaluate the guardrails themselves.

  5. 5

    Monitor an LLM feature in production with online evals, dashboards and alerts.

  6. 6

    Run an incident review for an AI failure and write it up.

Roles this prepares you for

  • AI quality engineer
  • ML reliability engineer
  • AI engineer (evals)
  • Applied scientist (evaluation)

No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.

Your record at week 12.

Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.

Graded, line by line
Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
Six gates
A block does not open until the previous project passes its gate. You always know where you are and what is next.
A weekly 1-1
Thirty minutes with your instructor, who has already read your code before the call.
One page an employer can run
Every grade, project and the viva recording at /verify. An employer opens it, runs the code and watches you defend it.

Record, LLM Evals and Reliability Bootcamp

Example

  1. Week 2

    Golden set

    Two labellers agree above a threshold on the set

    Graded, gate signed
  2. Week 4

    Calibrated judge

    Judge agreement with humans is measured and reported

    Graded, gate signed
  3. Week 6

    Regression in CI

    A seeded regression is caught before merge

    Graded, gate signed
  4. Week 8

    Guardrails

    The attack suite is blocked without hurting the golden set

    Graded, gate signed
  5. Week 10

    Monitoring and incident

    A viva: defend the monitoring and walk through the incident

    Graded, gate signed
  6. Week 12

    Employer brief

    The brief's owner accepts the result

    Graded, gate signed
  7. Weeks 1 to 12

    Twelve 1-1 notes

    What your instructor saw in your work each week, and what you agreed to do next.

    Kept
  8. Week 12

    Your hiring plan and its outcome

    Target roles, the gap map, applications, interviews and where you landed.

    Kept

The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.

8 entries, twelve weeks. One email when applications open.

No account, no card. One email when it opens; we never sell before it exists.

How we help you find a job.

The last block is not curriculum. It is the hiring sprint, and the proof you built in the ten weeks before it.

  1. 1

    Your hiring plan, from week 1

    Six parts you and your instructor keep: target roles, the gap map from real postings, the proof to send, weekly actions, an interview log, the outcome. Read before every 1-1.

  2. 2

    The hiring sprint

    Weeks 11 and 12: CV, portfolio, applications, mock interviews, and demo day in front of hiring partners.

  3. 3

    A record an employer can run

    Your six deployed projects and the recorded viva on /verify. An employer opens it, runs the code and watches you defend it.

  4. 4

    The career agent

    Paste a real job posting at /career and it maps the role to your graded work and what to do next.

  5. 5

    The roles directory

    Every role we prepare people for, what it pays and what it asks, at /roles.

One email when applications open.

Fifty seats, one instructor, twelve weeks. The waitlist hears the date and the price first, and nothing is charged before the cohort exists.

No account, no card. One email when it opens; we never sell before it exists.