Yes, but as an engineering skill, not a job title: the standalone "prompt engineer" role has largely faded, and the work has moved inside ordinary software engineering, where it means writing system prompts, defining structured output, choosing examples, describing tools so a model calls them correctly, and running an eval before you ship a prompt change. Employers now ask for the practice around prompts more than the name. In 315 AI and machine learning job ads we collected on 19 September 2026, 17 (5%) mentioned prompt engineering, while 73 (23%) mentioned evals and 123 (39%) mentioned agents.
So the honest answer for a developer is: learn it, and learn it as part of building software rather than as a career of its own.
What happened to the prompt engineer job title?
Early on, getting a useful answer out of a model took tricks: careful phrasing and long lists of reminders. Models have become much better at following plain instructions, so much of that craft stopped mattering. The prompt is now one component of a system that has inputs, outputs, tools, tests and a version history, and teams that ship LLM features rarely hire someone only to write the words. They expect the engineer who owns the feature to own its prompts too.
The job ads show the shift. Prompt engineering is named in 17 of the 315 postings. The terms that appear far more often describe what prompts are now for: agents (123 postings), evals (73), and the model providers' APIs themselves, with Anthropic or Claude named in 49 (16%) and the OpenAI API in 28 (9%). A developer who can make an agent call the right tool, and prove that a change did not break it, is doing prompt engineering whether the ad uses the phrase or not.
The small print: the sample is Hacker News "Who is hiring?" threads for July to September 2026, Remotive and Arbeitnow. It leans towards startups and Europe, it is a keyword count rather than a study, and a posting can mention a term without making it central. The method and the full counts are in what AI job ads ask for, and the data is public.
What does prompt engineering look like inside engineering work now?
Five jobs, each with a way to check it worked:
| Task | What it involves | How you know it works |
|---|---|---|
| System prompt | The role, the rules, what to refuse, the output format; kept in version control like code | An eval on real inputs, before and after each change |
| Structured output | A schema the model's answer must match, validated in your code, with a retry or a clean failure | The share of held-out inputs that validate |
| Examples | A few worked examples chosen to cover the awkward cases, not only the easy one | The eval score with and without them |
| Tool descriptions | The name, description and parameter schema the model reads to decide when to call a tool | A test set of requests with the tool call you expect |
| Evals | A golden set of inputs with expected outputs, and a scorer, run on every change | A number you can compare across versions |
Tool descriptions are the part developers most often forget is a prompt: the model chooses a tool from its description, so a vague one produces wrong calls. Our explainer on tool use and function calling covers the mechanics.
Structured output is where prompting meets ordinary code. You define the shape you need and refuse anything else:
from pydantic import BaseModel, ValidationError
class Invoice(BaseModel):
supplier: str
total: float
currency: str
def parse(raw: str) -> Invoice | None:
try:
return Invoice.model_validate_json(raw)
except ValidationError:
return None # log it, retry once, or send it to a person
The prompt asks for the format; the validator guarantees it. Neither is enough alone.
Why do evals matter more than clever wording?
Because a prompt change is a code change with invisible side effects. You fix the case a user complained about and quietly break three others, and nobody notices until a customer does. An eval makes the side effects visible: a set of real inputs, the answer you expect for each, and a score you run before and after every edit. That is also why employers name evals far more often than prompt engineering: the eval is the part that makes prompting trustworthy enough to ship. Our guide to evaluating LLM outputs explains golden sets, rubrics and when to use a model as a judge. Start small. A few dozen real inputs with checked answers will catch more regressions than any amount of careful rewording.
When is prompting not enough?
When the model lacks facts, or when the behaviour you need will not hold however you phrase it. Missing or changing facts, such as your documents, policies or prices, are a retrieval problem: put the facts in front of the model at question time. A behaviour that must be consistent at high volume, in a narrow task, may justify fine-tuning, but only after you have measured the best prompt and found it short. Our comparison of fine-tuning and prompt engineering sets out where each wins, and should you fine-tune an open-weights model covers the decision in detail. In both cases, the eval you built for your prompt is what tells you whether the heavier option helped.
Will prompt engineering disappear completely?
Not while models take instructions in language. What fades is the wording craft: models keep getting better at inferring what you meant, so tricks matter less each year. What stays is specification and measurement. Someone still has to decide what the system should do, what format it returns, what it must refuse, which tool it should reach for, and how to tell whether it is doing all of that. As agents take more actions on their own, that specification work matters more, not less, because a vague instruction now produces a wrong action rather than a wrong paragraph.
If you are not a developer, prompting well is still a useful everyday skill, but it is a different subject from the one this piece covers. The free generative AI skill check takes about three minutes and shows where you stand.
Where should you start this week?
- Take one prompt in your codebase, or write one for a small real task, and move it into its own file under version control.
- Define a schema for its output and validate every response in code.
- Collect a few dozen real inputs, write down the answer you expect for each, and score the current prompt.
- Make one change, score again, and keep or revert it on the number.
- If the prompt uses tools, write a set of requests with the tool call you expect for each, and check the model makes them.
Where does Square 1 teach this?
Prompt Engineering for Developers is an on-demand course recorded by an instructor and graded by Nova, the AI tutor: about eight hours across six modules, for developers who work in Python or TypeScript. The graded work is a prompt that passes a schema check, a tool-calling prompt with validation, a prompt eval with a golden set, a versioned prompt library in a repository, and a final production prompt with its eval and history. Building with the Claude API takes the same skills into a shipped application with tool use, streaming and evals. The Gen AI Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, in six blocks that run from a deployed assistant with checked structured output to a production LLM product with monitoring, and it asks for Python and Git. All three are taking a waitlist today. The AI engineer role page describes where these skills are used day to day.
