Skip to content
Hiring & Assessment

How do you get a job in AI evaluation (evals) in 2026?

Evals are a skill employers hire for inside other roles, not yet a job title: 73 of 315 AI job ads (23%) asked for evals, but only 2 of 3,610 postings had evaluation in the title. Which roles ask for it most, the skills involved, and how to prove you can do it.

Nikhil De Silva · Founder, Square 1 AI6 min read

You get a job in AI evaluation by becoming the engineer on an AI team who can prove whether the system works, because in 2026 evals are a skill employers hire for inside other roles, not yet a job title of their own. In our count of 315 AI job ads from July to September 2026, 73 (23%) asked for evals or evaluation. In a wider re-run of 3,610 postings on 29 September, only 2 had evaluation in the title. The demand is real; it just arrives under the name AI engineer, forward deployed engineer or product manager.

That changes how you should go about it. Searching for "evals engineer" will find almost nothing. Building the evidence that you can do evals, then applying for the roles that list them, is the route.

What is an AI evaluation job, actually?

Evaluation work answers one question with numbers: how good is this AI system, on the cases that matter, and did the last change make it better or worse? In practice that means:

  • Defining quality. Turning "the answers should be good" into a rubric two people apply the same way.
  • Golden sets. Collecting and labelling real cases, including the awkward ones, so there is something to test against.
  • Judges. Using a model to score outputs at scale, and checking that its scores agree with human labels before anyone trusts them.
  • Regression testing. Running the eval suite on every prompt or model change, so a bad change is caught before it ships.
  • Production monitoring. Sampling live traffic, scoring it and alerting when quality drifts.
  • Incident review. When the system does something wrong in public, working out why and adding the case to the suite.

Our explainer on how to evaluate LLM outputs covers the methods in more depth.

How many jobs ask for evals?

More than most people expect, and in more roles. From our headline sample of 19 September and the 29 September re-runs:

Postings Asking for evals
All AI and ML ads (19 Sep) 73 of 315 (23%)
Forward deployed engineer 9 of 29 (31%)
Sales or solutions engineer 10 of 49 (20%)
Software engineering ads that name AI coding tools 11 of 67 (16%)
Product manager 12 of 109 (11%)
Titled as an evals or evaluation role 2 postings in 3,610

For scale, in the same 315 AI ads Python appeared in 38% and agents in 39%. Evals at 23% sit not far behind the core stack. The forward deployed figure is the telling one: engineers who deploy AI inside a customer's organisation are asked for evals more often than AI ads overall, because the customer will not sign off on a demo. They sign off on "it handles these cases, and here are the ones it does not".

The small print: these are public postings from Hacker News "Who is hiring?" (July to September 2026), Remotive and Arbeitnow, so they lean towards startups and Europe, and Australia barely appears. Some Arbeitnow postings are duplicated, so read these as postings, not companies. The role-level counts are small (29 FDE ads, 2 evals-titled ads), and the files are public: headline, by role and more roles.

Why is "evals engineer" not a common title yet?

Because most teams are not big enough to have one person do only evaluation. On a team of five building an AI product, evals are one of several things the AI engineer owns, alongside retrieval, tool use and deployment. The two evals-titled postings we found are the exception, and neither stated a number of years, so there is no pattern to read from them.

This is normal for a skill that is becoming a discipline. Testing went the same way: first something every developer was meant to do, then a specialism at larger companies. For now, the practical conclusion is that evals make you a stronger candidate for a job with a broader title, and the person who owns quality on a small team often becomes the obvious hire when a larger team creates the dedicated role.

What skills do you need for evaluation work?

Five, and the last two are where most candidates are thin:

  1. Python, enough to build a harness, call models and process results.
  2. Enough statistics to not fool yourself. Sample sizes, variance between runs, agreement between labellers. An eval whose score swings between identical runs tells you nothing until you know by how much.
  3. Rubric writing and labelling. Harder than it sounds: most disagreements about AI quality are really disagreements about what "good" means.
  4. Calibrated judges. Knowing that an LLM judge is only useful once you have measured its agreement with humans, and knowing how it drifts.
  5. Evals in the build. Suites that run in CI and block a bad change, plus monitoring once the system is live.

Adjacent to this is adversarial testing: checking that a system resists prompt injection and misuse, not only that it answers well. Red-teaming is its own path, but the two share the same habit of trying to break things with evidence.

Our AI engineer role page lists where evals sit in the wider job.

How do you prove you can do evals?

With one feature, evaluated properly, in a public repository. Pick a narrow LLM feature, say answering support questions from a help centre. Then:

  • Write a rubric and label a golden set of real cases, with a note on where you and a second labeller disagreed.
  • Build an LLM judge and report how often it agrees with your labels.
  • Put the suite in CI and show a deliberately bad prompt change being blocked.
  • Add a short results report: what the feature gets right, what it gets wrong, and what you would fix first.

That is four artefacts a hiring manager can open, and it answers the question they have about every candidate who says they "know evals".

Where do you start this week?

Take any LLM feature you have built or can build in an afternoon. Write 30 test cases by hand, including five you expect it to fail. Score its outputs yourself against a three-point rubric, then ask a friend to score the same outputs without seeing yours, and count where you disagree. That disagreement is the first real lesson in evaluation, and most people who say they do evals have never measured it.

Which Square 1 programme fits?

The LLM Evals and Reliability Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week. Its six blocks each end in a deployed project and a gate: a labelled golden set with measured labeller agreement, a calibrated judge deployed as a service, an eval pipeline in CI that catches a seeded regression, a guardrail layer with its own evals, a monitored LLM feature with a written incident review, and an employer brief in the final block alongside a hiring sprint. The recorded viva asks you to defend the monitoring and walk through the incident. The entry bar is having shipped an LLM feature. Evaluating AI Systems is a shorter on-demand course, recorded by an instructor and graded by Nova, the AI tutor, covering rubrics, golden sets, calibrated judges and CI suites. If the adversarial side interests you more, the AI Security and Red-Teaming Bootcamp covers attacking and defending LLM applications. All are taking waitlist places today.

Questions people ask

Is AI evaluation a real job?

The work is; the title is rare. In 315 AI job ads from July to September 2026, 73 (23%) asked for evals or evaluation, but in 3,610 postings collected on 29 September only 2 had evaluation in the job title. Most evaluation work is done by AI engineers, forward deployed engineers and product teams.

Which jobs ask for evals skills?

Forward deployed engineer ads asked for evals most often in our sample (9 of 29, 31%), then sales or solutions engineers (10 of 49), software engineering ads that name AI coding tools (11 of 67) and product managers (12 of 109). Across all 315 AI ads the figure was 23%.

What does an AI evaluation engineer do?

Defines quality as a rubric people agree on, builds labelled golden sets, calibrates LLM judges against human labels, runs eval suites in CI so bad changes are blocked, monitors quality in production and reviews incidents when the system fails.

What skills do you need to work in LLM evaluation?

Python, enough statistics to handle sample sizes and run-to-run variance, rubric writing and labelling, LLM judges calibrated against human labels, and eval suites that run in CI and in production monitoring.

How do you show an employer you can do evals?

Evaluate one LLM feature properly in a public repository: a rubric and labelled golden set with labeller agreement, an LLM judge with its agreement to human labels reported, a CI suite that blocks a bad prompt change, and a short results report.

Free skill check · about 3 minutes

Where do you stand on Generative AI?

Five questions, and a skill breakdown the moment you finish: your strengths, the gaps to close, and what to learn next from real curriculum.

Start the Generative AI skill check

Free, with a student account — the check is the first entry in your record.

Learn this by building it

The programmes that teach what this piece covers, each ending in deployed work graded against a rubric you can read.