You get a job in AI evaluation by becoming the engineer on an AI team who can prove whether the system works, because in 2026 evals are a skill employers hire for inside other roles, not yet a job title of their own. In our count of 315 AI job ads from July to September 2026, 73 (23%) asked for evals or evaluation. In a wider re-run of 3,610 postings on 29 September, only 2 had evaluation in the title. The demand is real; it just arrives under the name AI engineer, forward deployed engineer or product manager.
That changes how you should go about it. Searching for "evals engineer" will find almost nothing. Building the evidence that you can do evals, then applying for the roles that list them, is the route.
What is an AI evaluation job, actually?
Evaluation work answers one question with numbers: how good is this AI system, on the cases that matter, and did the last change make it better or worse? In practice that means:
- Defining quality. Turning "the answers should be good" into a rubric two people apply the same way.
- Golden sets. Collecting and labelling real cases, including the awkward ones, so there is something to test against.
- Judges. Using a model to score outputs at scale, and checking that its scores agree with human labels before anyone trusts them.
- Regression testing. Running the eval suite on every prompt or model change, so a bad change is caught before it ships.
- Production monitoring. Sampling live traffic, scoring it and alerting when quality drifts.
- Incident review. When the system does something wrong in public, working out why and adding the case to the suite.
Our explainer on how to evaluate LLM outputs covers the methods in more depth.
How many jobs ask for evals?
More than most people expect, and in more roles. From our headline sample of 19 September and the 29 September re-runs:
| Postings | Asking for evals |
|---|---|
| All AI and ML ads (19 Sep) | 73 of 315 (23%) |
| Forward deployed engineer | 9 of 29 (31%) |
| Sales or solutions engineer | 10 of 49 (20%) |
| Software engineering ads that name AI coding tools | 11 of 67 (16%) |
| Product manager | 12 of 109 (11%) |
| Titled as an evals or evaluation role | 2 postings in 3,610 |
For scale, in the same 315 AI ads Python appeared in 38% and agents in 39%. Evals at 23% sit not far behind the core stack. The forward deployed figure is the telling one: engineers who deploy AI inside a customer's organisation are asked for evals more often than AI ads overall, because the customer will not sign off on a demo. They sign off on "it handles these cases, and here are the ones it does not".
The small print: these are public postings from Hacker News "Who is hiring?" (July to September 2026), Remotive and Arbeitnow, so they lean towards startups and Europe, and Australia barely appears. Some Arbeitnow postings are duplicated, so read these as postings, not companies. The role-level counts are small (29 FDE ads, 2 evals-titled ads), and the files are public: headline, by role and more roles.
Why is "evals engineer" not a common title yet?
Because most teams are not big enough to have one person do only evaluation. On a team of five building an AI product, evals are one of several things the AI engineer owns, alongside retrieval, tool use and deployment. The two evals-titled postings we found are the exception, and neither stated a number of years, so there is no pattern to read from them.
This is normal for a skill that is becoming a discipline. Testing went the same way: first something every developer was meant to do, then a specialism at larger companies. For now, the practical conclusion is that evals make you a stronger candidate for a job with a broader title, and the person who owns quality on a small team often becomes the obvious hire when a larger team creates the dedicated role.
What skills do you need for evaluation work?
Five, and the last two are where most candidates are thin:
- Python, enough to build a harness, call models and process results.
- Enough statistics to not fool yourself. Sample sizes, variance between runs, agreement between labellers. An eval whose score swings between identical runs tells you nothing until you know by how much.
- Rubric writing and labelling. Harder than it sounds: most disagreements about AI quality are really disagreements about what "good" means.
- Calibrated judges. Knowing that an LLM judge is only useful once you have measured its agreement with humans, and knowing how it drifts.
- Evals in the build. Suites that run in CI and block a bad change, plus monitoring once the system is live.
Adjacent to this is adversarial testing: checking that a system resists prompt injection and misuse, not only that it answers well. Red-teaming is its own path, but the two share the same habit of trying to break things with evidence.
Our AI engineer role page lists where evals sit in the wider job.
How do you prove you can do evals?
With one feature, evaluated properly, in a public repository. Pick a narrow LLM feature, say answering support questions from a help centre. Then:
- Write a rubric and label a golden set of real cases, with a note on where you and a second labeller disagreed.
- Build an LLM judge and report how often it agrees with your labels.
- Put the suite in CI and show a deliberately bad prompt change being blocked.
- Add a short results report: what the feature gets right, what it gets wrong, and what you would fix first.
That is four artefacts a hiring manager can open, and it answers the question they have about every candidate who says they "know evals".
Where do you start this week?
Take any LLM feature you have built or can build in an afternoon. Write 30 test cases by hand, including five you expect it to fail. Score its outputs yourself against a three-point rubric, then ask a friend to score the same outputs without seeing yours, and count where you disagree. That disagreement is the first real lesson in evaluation, and most people who say they do evals have never measured it.
Which Square 1 programme fits?
The LLM Evals and Reliability Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week. Its six blocks each end in a deployed project and a gate: a labelled golden set with measured labeller agreement, a calibrated judge deployed as a service, an eval pipeline in CI that catches a seeded regression, a guardrail layer with its own evals, a monitored LLM feature with a written incident review, and an employer brief in the final block alongside a hiring sprint. The recorded viva asks you to defend the monitoring and walk through the incident. The entry bar is having shipped an LLM feature. Evaluating AI Systems is a shorter on-demand course, recorded by an instructor and graded by Nova, the AI tutor, covering rubrics, golden sets, calibrated judges and CI suites. If the adversarial side interests you more, the AI Security and Red-Teaming Bootcamp covers attacking and defending LLM applications. All are taking waitlist places today.
