Courses / On-demand courses
Running AI on Kubernetes
Serve, scale and pay for models on Kubernetes without the bill or the latency surprising you.
Running AI on Kubernetes is a self-paced online course: 9.5 hours of recorded sessions across 7 modules, with 6 graded by Nova, Square 1's AI tutor. The founding price is A$45, down from A$90. Start any time and keep your own pace.
Price: A$45A$9050% off
Founding price for the first cohort, until 21 October.
Founding price, one payment. Founding students keep their founding price on any second programme.
Founding intake
Spot holders hear the start date first.
Module by module
Module 1 · 60 min
What changes when the workload is a model
Module 2 · 90 min · graded
Serving a model in a container
Module 3 · 90 min · graded
GPUs, node pools and scheduling
Module 4 · 90 min · graded
Autoscaling on real load
Module 5 · 60 min · graded
Batch jobs and pipelines
Module 6 · 60 min · graded
Monitoring latency, errors and cost
Module 7 · 120 min · graded
Project: an LLM service in production shape
- recorded hours
- 9.5
- modules, in order
- 7
- deployed projects
- 3
- graded checkpoints
- 6
Who it is for, and what you receive.
Engineers who know Kubernetes basics and some Python and want to serve, scale and pay for AI models on a cluster without surprises.
Before you start
Kubernetes basics (Kubernetes from Zero, or the same experience) and some Python.
Skills
- Kubernetes
- model serving
- GPUs
- autoscaling
- observability
- cost control
- 1
The recorded sessions
Taught by an instructor, taken in order at your own pace, yours for twelve months.
- 2
Graded work after every module
Exercises and projects marked by Nova against a rubric you can read, with what you did well and what to fix.
- 3
The projects
Deployed and graded, each one on your record with its repository.
- 4
Your hiring plan
The same six-part plan the bootcamps use, run with the career agent.
- 5
The record and the certificate
Every grade at square1ai.com/u/{handle}; a credential ID that resolves at square1ai.com/verify.
- 6
Nova as your tutor
Help that is about your actual work, because it has read all of it.
The content plan, module by module.
9.5 recorded hours from an instructor, taken in order at your own pace. Nova grades the work at the end of each module.
What changes when the workload is a model
Module 1 · 60 min
Watch and take notes; nothing to submit in this module.
- Kubernetes
- Docker
- Python
Memory and start-up time
GPUs and their cost
Latency budgets
Serving a model in a container
Module 2 · 90 min
Graded at the end: A model served behind an API on the cluster.
- Docker
- Kubernetes
- Python
Packaging a model
Serving runtimes
An API in front
GPUs, node pools and scheduling
Module 3 · 90 min
Graded at the end: A GPU node pool that only model pods use.
- TypeScript
- Kubernetes
- Docker
GPU node pools
Taints, tolerations and affinity
Sharing a GPU
Autoscaling on real load
Module 4 · 90 min
Graded at the end: An autoscaler tuned against a load test.
- Kubernetes
- Docker
- Python
Load testing
Scaling on custom metrics
Scale to zero
Batch jobs and pipelines
Module 5 · 60 min
Graded at the end: A scheduled batch job with retries.
- Kubernetes
- Docker
- Python
Jobs and CronJobs
Retries and idempotency
Queues
Monitoring latency, errors and cost
Module 6 · 60 min
Graded at the end: A dashboard with a cost-per-request line.
- Kubernetes
- Docker
- Python
Metrics and traces
Dashboards
Cost per request
Project: an LLM service in production shape
Module 7 · 120 min
Graded at the end: The service under load, with its cost report.
- Kubernetes
- Docker
- Python
The service
The load test
The dashboard
The cost report
The projects.
Each is deployed and graded by Nova against a rubric you can read before you start.
- 1
A model served on Kubernetes
Package an open-weights model or a small classifier in a container and serve it behind an API on a cluster, with health checks and resource limits that match what it actually uses.
You hand in
- Serving image
- Manifests
- API endpoint
- Resource measurements
The rubric requires
The API answers under a stated latency with the declared resource limits.
- 2
Autoscaling under load
Load test your served model, then tune autoscaling so it keeps latency under budget at peak and scales back down when traffic falls. Show the before and after.
You hand in
- Load test scripts
- Autoscaler configuration
- Before-and-after latency and replica charts
The rubric requires
Latency stays under the budget at the tested peak and replicas fall back after it.
- 3
An LLM service in production shape
Run an LLM-backed service on a cluster with monitoring for latency, errors and cost, and write a one-page cost report a manager could act on.
You hand in
- Service repository
- Dashboard
- Cost-per-request report
- Runbook
The rubric requires
The dashboard shows live latency, errors and cost per request, and the report's numbers match it.
What you can do at the end.
AI services you have run and costed on Kubernetes, and three graded projects on your record that show it.
- 1
Explain what changes when a Kubernetes workload is a model.
- 2
Serve a model in a container behind an API on a cluster.
- 3
Schedule GPU workloads onto the right nodes and keep other pods off them.
- 4
Tune autoscaling against a measured load test.
- 5
Run batch inference as scheduled jobs with retries.
- 6
Monitor latency, errors and cost per request.
Roles this prepares you for
- MLOps engineer
- Platform engineer
- Cloud engineer
No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.
Your record at the end.
Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.
- Graded, line by line
- Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
- Module by module
- Each module ends in graded work; the next opens when you are ready, on your own schedule.
- Nova remembers
- Help is about your actual work, because the tutor has every submission and every failed exercise of yours.
- One page an employer can run
- Every grade and project at /verify. An employer opens it and runs the code.
Record, Running AI on Kubernetes
Example
- ProjectGraded
A model served on Kubernetes
The API answers under a stated latency with the declared resource limits.
- ProjectGraded
Autoscaling under load
Latency stays under the budget at the tested peak and replicas fall back after it.
- ProjectGraded
An LLM service in production shape
The dashboard shows live latency, errors and cost per request, and the report's numbers match it.
- Every moduleGraded
6 graded checkpoints
Each module ends in work Nova grades line by line against a rubric you can read.
- AfterKept
Your hiring plan
Target roles, the gap map, the proof to send, weekly actions and an interview log.
The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.
5 entries, at your own pace. 25 spots are open.
Reserve my spotHow we help you find a job.
Proof, not a certificate: the projects you deployed are the thing you show, and the tools below are yours to use.
- 1
Your hiring plan
The same six-part plan the bootcamps use: target roles, the gap map, the proof to send, weekly actions, an interview log, the outcome. You run it with the career agent.
- 2
A record an employer can run
Your graded projects on /verify. An employer opens it and runs the code.
- 3
The career agent
Paste a real job posting at /career and it maps the role to your graded work and what to do next.
- 4
The roles directory
Every role we prepare people for, what it pays and what it asks, at /roles.
- 5
A path to the live cohort
If you want the instructor, the gates and the hiring sprint, the bootcamp on the same subject is one waitlist away.
Questions people ask.
About Running AI on Kubernetes, answered from the plan on this page.
Related programmes
- Kubernetes from Zero · 10 h on demand
- Cloud Engineer with AI Bootcamp · 12-week live bootcamp
- Cloud Infrastructure with Terraform · 9.5 h on demand
How long is Running AI on Kubernetes?
9.5 hours of recorded sessions across 7 modules, taken at your own pace. 6 of the modules end in work Nova grades.
Is Running AI on Kubernetes live or self-paced?
Self-paced. The sessions are recorded by an instructor who does this work, and Nova, Square 1's AI tutor, grades every checkpoint and project against a rubric you can read.
How much does Running AI on Kubernetes cost?
A$45 for the founding intake, paid once; the standard price is A$90. It is the same price in every country. Founding students keep their founding price on any second programme.
Who is Running AI on Kubernetes for?
Engineers who know Kubernetes basics and some Python and want to serve, scale and pay for AI models on a cluster without surprises. Before you start: Kubernetes basics (Kubernetes from Zero, or the same experience) and some Python.
What will I build in Running AI on Kubernetes?
3 deployed projects: A model served on Kubernetes, Autoscaling under load and An LLM service in production shape. Each is graded against a published rubric and kept on a record an employer can open at /verify.
What skills does Running AI on Kubernetes teach?
Kubernetes, model serving, GPUs, autoscaling, observability and cost control. It prepares you for roles such as MLOps engineer, Platform engineer and Cloud engineer.
Does Running AI on Kubernetes guarantee a job?
No. No placement rate is published because there are no graduates to count yet. What you leave with is graded, deployed work on one record an employer can open and run, and a hiring plan you keep with the career agent.
How do I join Running AI on Kubernetes?
Reserve one of the twenty-five founding spots on this page with your email. Spot holders hear the opening date first, and nothing is charged before you confirm. The price on this page is in Australian dollars and is the same in every country.
25 spots. Hold one of them.
25 founding spots. Recorded by an instructor who does this work, graded by Nova. Spot holders hear the opening date first, and nothing is charged before you confirm.
- You can run AI workloads on Kubernetes and explain their cost
- Three graded projects on your record
About a minute. No account, no card, and nothing is charged until you confirm.
