Courses / Live bootcamps
LLM Evals and Reliability Bootcamp
Make an AI system you can trust: evals, guardrails, monitoring, incident review.
- weeks, live on Zoom
- 12
- hours a week
- 15
- deployed projects
- 6
- minutes 1-1, every week
- 30
Who it is for, and what you receive.
Engineers who have shipped an LLM feature and been burned by it, and teams that need someone to own quality.
Before you start
You have shipped an LLM feature and been burned by it.
Skills
- Python
- evals
- LLM judges
- CI
- guardrails
- monitoring
- incident review
A week, about 15 hours
- Live on Zoom
- 4h
- Recorded lessons and exercises
- 6h
- The project
- 5h
- 1
Six deployed projects
Each graded by Nova against a published rubric and signed off by the instructor at the gate.
- 2
The recorded lessons
The track's lessons and exercises, graded line by line, open for twelve months after the cohort ends.
- 3
Live on Zoom every week
A 90-minute live code review, a 60-minute squad lab, 60 minutes of office hours, and a 30-minute 1-1 with your instructor.
- 4
Your hiring plan
Opened in week 1 and reviewed in every 1-1: target roles, the gap map, the proof list, weekly actions, an interview log, the outcome.
- 5
The record
Every grade and every project with its repository at square1ai.com/u/{handle}, and the recording of your viva, verifiable by any employer.
- 6
The certificate
A credential ID that resolves at square1ai.com/verify to the real completion.
- 7
The hiring sprint
Weeks 11 and 12: CV and portfolio from graded work, applications, mock interviews scored against the role, demo day.
- 8
Nova for twelve weeks
A tutor with every submission of yours in its memory, at 2am as well as in class.
Twelve weeks in six blocks.
Each block teaches for a week, then you build and deploy a project, then a gate checks it before the next block opens. The project and its gate are the block's record entry.
Live time each week on Zoom: a 90-minute class where the instructor reviews real submissions, a 60-minute squad lab, 60 minutes of office hours, and your own 30-minute 1-1. Nothing is lectured live; the recorded lessons do that.
- 1
What good looks like
Weeks 1 to 2Week 1
What good looks like
- Defining quality
- Rubrics and inter-rater agreement
- Sampling and labelling
- Golden sets
You build: A rubric and a labelling plan
Week 2
Project 1: the golden set
- Labelling 300 cases
- Agreement reports
You build: Labelled golden set, published
Project 1, the gate at week 2
Golden set
Define quality for a support-answer feature and label 300 cases.
You hand in
- Rubric
- Labelling guide
- Golden set
- Agreement report
The gate
Two labellers agree above the stated threshold.
- 2
Judges
Weeks 3 to 4Week 3
Judges
- LLM-as-judge
- Pairwise and absolute scoring
- Calibration
- Judge drift
You build: A calibrated judge
Week 4
Project 2: the judge
- Judge as a service
- Measuring agreement
You build: Judge service with measured agreement
Project 2, the gate at week 4
Calibrated judge
A judge that scores answers the way the labellers do.
You hand in
- Judge service
- Calibration report
- Agreement with humans
The gate
Judge agreement measured and reported.
- 3
Regression and CI
Weeks 5 to 6Week 5
Regression and CI
- Eval suites in the build
- Prompt and model versioning
- Flakiness
- Sampling budgets
You build: Eval pipeline; squads form
Week 6
Project 3: the pipeline
- Seeded regressions
- Change control
You build: CI pipeline blocking a bad prompt change
Project 3, the gate at week 6
Regression in CI
Make prompt and model changes safe.
You hand in
- Eval pipeline
- Version control for prompts
- Sampling policy
The gate
A seeded regression is caught before merge.
- 4
Guardrails
Weeks 7 to 8Week 7
Guardrails
- Input and output checks
- Prompt injection defences
- PII and policy
- Evaluating the guardrails
You build: Guardrail layer
Week 8
Project 4: the guardrails
- Attack suites
- Protecting the golden set
You build: Guardrails deployed with their own evals
Project 4, the gate at week 8
Guardrails
Stop the attacks without hurting the good cases.
You hand in
- Input and output guards
- Injection suite
- PII policy
- Guardrail evals
The gate
The attack suite is blocked; the golden set is unharmed.
- 5
Production monitoring
Weeks 9 to 10Week 9
Production monitoring
- Online evals and sampling
- Dashboards that mean something
- Alerting
- Incident review for AI
You build: Monitoring design
Week 10
Project 5 and the viva
- The monitored feature
- Walking through an incident
You build: Monitored feature with alerts and a written incident review; the viva
Project 5, the gate at week 10
Monitoring and incident
Run the feature in production and survive an incident.
You hand in
- Online evals
- Dashboard
- Alerts
- A written incident review
The gate
The viva: defend the monitoring and walk through the incident.
- 6
Employer brief and hiring sprint
Weeks 11 to 12Week 11
Employer brief
- A real problem from a hiring partner, in squads
- A partner's reliability problem
- Working to someone else's definition of done
You build: The employer brief in progress
Week 12
Hiring sprint
- CV and portfolio built from graded work
- Applications and follow-ups in the hiring plan
- Mock interviews scored against the role
- Demo day
You build: Project 6 delivered; demo day
Project 6, the gate at week 12
Employer brief
A partner's reliability problem.
You hand in
- The brief delivered
- Demo day
The gate
The brief's owner accepts the result.
What you can do by week 12.
The quality infrastructure every AI team needs and few have: a golden set, a calibrated judge, a CI pipeline, guardrails and monitoring, built once and carried to any team.
- 1
Define quality for an LLM feature as a rubric people agree on, and build a labelled golden set.
- 2
Build LLM judges calibrated against human labels and measure their agreement.
- 3
Run eval suites in CI with change control for prompts and models, and handle flakiness and sampling.
- 4
Build guardrails for inputs, outputs, prompt injection and PII, and evaluate the guardrails themselves.
- 5
Monitor an LLM feature in production with online evals, dashboards and alerts.
- 6
Run an incident review for an AI failure and write it up.
Roles this prepares you for
- AI quality engineer
- ML reliability engineer
- AI engineer (evals)
- Applied scientist (evaluation)
No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.
Your record at week 12.
Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.
- Graded, line by line
- Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
- Six gates
- A block does not open until the previous project passes its gate. You always know where you are and what is next.
- A weekly 1-1
- Thirty minutes with your instructor, who has already read your code before the call.
- One page an employer can run
- Every grade, project and the viva recording at /verify. An employer opens it, runs the code and watches you defend it.
Record, LLM Evals and Reliability Bootcamp
Example
- Week 2Graded, gate signed
Golden set
Two labellers agree above a threshold on the set
- Week 4Graded, gate signed
Calibrated judge
Judge agreement with humans is measured and reported
- Week 6Graded, gate signed
Regression in CI
A seeded regression is caught before merge
- Week 8Graded, gate signed
Guardrails
The attack suite is blocked without hurting the golden set
- Week 10Graded, gate signed
Monitoring and incident
A viva: defend the monitoring and walk through the incident
- Week 12Graded, gate signed
Employer brief
The brief's owner accepts the result
- Weeks 1 to 12Kept
Twelve 1-1 notes
What your instructor saw in your work each week, and what you agreed to do next.
- Week 12Kept
Your hiring plan and its outcome
Target roles, the gap map, applications, interviews and where you landed.
The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.
8 entries, twelve weeks. One email when applications open.
How we help you find a job.
The last block is not curriculum. It is the hiring sprint, and the proof you built in the ten weeks before it.
- 1
Your hiring plan, from week 1
Six parts you and your instructor keep: target roles, the gap map from real postings, the proof to send, weekly actions, an interview log, the outcome. Read before every 1-1.
- 2
The hiring sprint
Weeks 11 and 12: CV, portfolio, applications, mock interviews, and demo day in front of hiring partners.
- 3
A record an employer can run
Your six deployed projects and the recorded viva on /verify. An employer opens it, runs the code and watches you defend it.
- 4
The career agent
Paste a real job posting at /career and it maps the role to your graded work and what to do next.
- 5
The roles directory
Every role we prepare people for, what it pays and what it asks, at /roles.
One email when applications open.
Fifty seats, one instructor, twelve weeks. The waitlist hears the date and the price first, and nothing is charged before the cohort exists.
