Skip to content

Courses / On-demand courses

On demandRecorded by an instructor, graded by Nova

Running AI on Kubernetes

Serve, scale and pay for models on Kubernetes without the bill or the latency surprising you.

Running AI on Kubernetes is a self-paced online course: 9.5 hours of recorded sessions across 7 modules, with 6 graded by Nova, Square 1's AI tutor. The founding price is A$45, down from A$90. Start any time and keep your own pace.

Price: A$45A$9050% off

Founding price for the first cohort, until 21 October.

Founding price, one payment. Founding students keep their founding price on any second programme.

Founding intake

Spot holders hear the start date first.

Reserve your spot

Three fields hold it. A few short questions after that.

No account, no card. Holding a spot is free, and you confirm before anything is charged. We use these details to place you, never to sell.

Get the course booklet

Running AI on Kubernetes. PDF, 11 pages, 260 KB. Tell us who you are and it downloads straight away.

You are a

No account, no card. We email you a copy and may contact you about this course. We never sell your details. Privacy

Module by module

  1. Module 1 · 60 min

    What changes when the workload is a model

  2. Module 2 · 90 min · graded

    Serving a model in a container

  3. Module 3 · 90 min · graded

    GPUs, node pools and scheduling

  4. Module 4 · 90 min · graded

    Autoscaling on real load

  5. Module 5 · 60 min · graded

    Batch jobs and pipelines

  6. Module 6 · 60 min · graded

    Monitoring latency, errors and cost

  7. Module 7 · 120 min · graded

    Project: an LLM service in production shape

recorded hours
9.5
modules, in order
7
deployed projects
3
graded checkpoints
6

Who it is for, and what you receive.

Engineers who know Kubernetes basics and some Python and want to serve, scale and pay for AI models on a cluster without surprises.

Before you start

Kubernetes basics (Kubernetes from Zero, or the same experience) and some Python.

Skills

  • Kubernetes
  • model serving
  • GPUs
  • autoscaling
  • observability
  • cost control
  1. 1

    The recorded sessions

    Taught by an instructor, taken in order at your own pace, yours for twelve months.

  2. 2

    Graded work after every module

    Exercises and projects marked by Nova against a rubric you can read, with what you did well and what to fix.

  3. 3

    The projects

    Deployed and graded, each one on your record with its repository.

  4. 4

    Your hiring plan

    The same six-part plan the bootcamps use, run with the career agent.

  5. 5

    The record and the certificate

    Every grade at square1ai.com/u/{handle}; a credential ID that resolves at square1ai.com/verify.

  6. 6

    Nova as your tutor

    Help that is about your actual work, because it has read all of it.

The content plan, module by module.

9.5 recorded hours from an instructor, taken in order at your own pace. Nova grades the work at the end of each module.

  1. What changes when the workload is a model

    Module 1 · 60 min

    Watch and take notes; nothing to submit in this module.

    • Kubernetes
    • Docker
    • Python

    Memory and start-up time

    GPUs and their cost

    Latency budgets

  2. Serving a model in a container

    Module 2 · 90 min

    Graded at the end: A model served behind an API on the cluster.

    • Docker
    • Kubernetes
    • Python

    Packaging a model

    Serving runtimes

    An API in front

  3. GPUs, node pools and scheduling

    Module 3 · 90 min

    Graded at the end: A GPU node pool that only model pods use.

    • TypeScript
    • Kubernetes
    • Docker

    GPU node pools

    Taints, tolerations and affinity

    Sharing a GPU

  4. Autoscaling on real load

    Module 4 · 90 min

    Graded at the end: An autoscaler tuned against a load test.

    • Kubernetes
    • Docker
    • Python

    Load testing

    Scaling on custom metrics

    Scale to zero

  5. Batch jobs and pipelines

    Module 5 · 60 min

    Graded at the end: A scheduled batch job with retries.

    • Kubernetes
    • Docker
    • Python

    Jobs and CronJobs

    Retries and idempotency

    Queues

  6. Monitoring latency, errors and cost

    Module 6 · 60 min

    Graded at the end: A dashboard with a cost-per-request line.

    • Kubernetes
    • Docker
    • Python

    Metrics and traces

    Dashboards

    Cost per request

  7. Project: an LLM service in production shape

    Module 7 · 120 min

    Graded at the end: The service under load, with its cost report.

    • Kubernetes
    • Docker
    • Python

    The service

    The load test

    The dashboard

    The cost report

The projects.

Each is deployed and graded by Nova against a rubric you can read before you start.

  1. 1

    A model served on Kubernetes

    Package an open-weights model or a small classifier in a container and serve it behind an API on a cluster, with health checks and resource limits that match what it actually uses.

    You hand in

    • Serving image
    • Manifests
    • API endpoint
    • Resource measurements

    The rubric requires

    The API answers under a stated latency with the declared resource limits.

  2. 2

    Autoscaling under load

    Load test your served model, then tune autoscaling so it keeps latency under budget at peak and scales back down when traffic falls. Show the before and after.

    You hand in

    • Load test scripts
    • Autoscaler configuration
    • Before-and-after latency and replica charts

    The rubric requires

    Latency stays under the budget at the tested peak and replicas fall back after it.

  3. 3

    An LLM service in production shape

    Run an LLM-backed service on a cluster with monitoring for latency, errors and cost, and write a one-page cost report a manager could act on.

    You hand in

    • Service repository
    • Dashboard
    • Cost-per-request report
    • Runbook

    The rubric requires

    The dashboard shows live latency, errors and cost per request, and the report's numbers match it.

What you can do at the end.

AI services you have run and costed on Kubernetes, and three graded projects on your record that show it.

  1. 1

    Explain what changes when a Kubernetes workload is a model.

  2. 2

    Serve a model in a container behind an API on a cluster.

  3. 3

    Schedule GPU workloads onto the right nodes and keep other pods off them.

  4. 4

    Tune autoscaling against a measured load test.

  5. 5

    Run batch inference as scheduled jobs with retries.

  6. 6

    Monitor latency, errors and cost per request.

Roles this prepares you for

  • MLOps engineer
  • Platform engineer
  • Cloud engineer

No placement rate is shown, because there are no graduates to count yet. The roles above are what the projects are built for.

Your record at the end.

Every exercise and project is graded by Nova against a rubric you can read, and every grade is kept on one page an employer can open and run. This is what the programme writes to it.

Graded, line by line
Nova reads every submission against the brief and the rubric and returns a score, what you did well and what to fix.
Module by module
Each module ends in graded work; the next opens when you are ready, on your own schedule.
Nova remembers
Help is about your actual work, because the tutor has every submission and every failed exercise of yours.
One page an employer can run
Every grade and project at /verify. An employer opens it and runs the code.

Record, Running AI on Kubernetes

Example

  1. Project

    A model served on Kubernetes

    The API answers under a stated latency with the declared resource limits.

    Graded
  2. Project

    Autoscaling under load

    Latency stays under the budget at the tested peak and replicas fall back after it.

    Graded
  3. Project

    An LLM service in production shape

    The dashboard shows live latency, errors and cost per request, and the report's numbers match it.

    Graded
  4. Every module

    6 graded checkpoints

    Each module ends in work Nova grades line by line against a rubric you can read.

    Graded
  1. After

    Your hiring plan

    Target roles, the gap map, the proof to send, weekly actions and an interview log.

    Kept

The entries, not the grades: those are yours to earn. The page lives at /verify and an employer needs no account to open it.

5 entries, at your own pace. 25 spots are open.

Reserve my spot

How we help you find a job.

Proof, not a certificate: the projects you deployed are the thing you show, and the tools below are yours to use.

  1. 1

    Your hiring plan

    The same six-part plan the bootcamps use: target roles, the gap map, the proof to send, weekly actions, an interview log, the outcome. You run it with the career agent.

  2. 2

    A record an employer can run

    Your graded projects on /verify. An employer opens it and runs the code.

  3. 3

    The career agent

    Paste a real job posting at /career and it maps the role to your graded work and what to do next.

  4. 4

    The roles directory

    Every role we prepare people for, what it pays and what it asks, at /roles.

  5. 5

    A path to the live cohort

    If you want the instructor, the gates and the hiring sprint, the bootcamp on the same subject is one waitlist away.

Get the course booklet

Running AI on Kubernetes. PDF, 11 pages, 260 KB. Tell us who you are and it downloads straight away.

You are a

No account, no card. We email you a copy and may contact you about this course. We never sell your details. Privacy

Questions people ask.

About Running AI on Kubernetes, answered from the plan on this page.

Related programmes

How long is Running AI on Kubernetes?

9.5 hours of recorded sessions across 7 modules, taken at your own pace. 6 of the modules end in work Nova grades.

Is Running AI on Kubernetes live or self-paced?

Self-paced. The sessions are recorded by an instructor who does this work, and Nova, Square 1's AI tutor, grades every checkpoint and project against a rubric you can read.

How much does Running AI on Kubernetes cost?

A$45 for the founding intake, paid once; the standard price is A$90. It is the same price in every country. Founding students keep their founding price on any second programme.

Who is Running AI on Kubernetes for?

Engineers who know Kubernetes basics and some Python and want to serve, scale and pay for AI models on a cluster without surprises. Before you start: Kubernetes basics (Kubernetes from Zero, or the same experience) and some Python.

What will I build in Running AI on Kubernetes?

3 deployed projects: A model served on Kubernetes, Autoscaling under load and An LLM service in production shape. Each is graded against a published rubric and kept on a record an employer can open at /verify.

What skills does Running AI on Kubernetes teach?

Kubernetes, model serving, GPUs, autoscaling, observability and cost control. It prepares you for roles such as MLOps engineer, Platform engineer and Cloud engineer.

Does Running AI on Kubernetes guarantee a job?

No. No placement rate is published because there are no graduates to count yet. What you leave with is graded, deployed work on one record an employer can open and run, and a hiring plan you keep with the career agent.

How do I join Running AI on Kubernetes?

Reserve one of the twenty-five founding spots on this page with your email. Spot holders hear the opening date first, and nothing is charged before you confirm. The price on this page is in Australian dollars and is the same in every country.

25 spots. Hold one of them.

25 founding spots. Recorded by an instructor who does this work, graded by Nova. Spot holders hear the opening date first, and nothing is charged before you confirm.

  • You can run AI workloads on Kubernetes and explain their cost
  • Three graded projects on your record
Reserve my spot

About a minute. No account, no card, and nothing is charged until you confirm.