You delegate a coding task to an agent like Codex without losing control by deciding what "done" means before it starts: a short written spec, a failing test or reproduction the agent must turn green, a sandbox that limits what it can touch, and a review of the diff that comes back as strict as any colleague's. The agent does the typing; you keep the definition of done and the decision to merge. Employers are starting to ask for this. Of 416 software engineering postings we collected on 29 September 2026, 67 (16%) named an AI coding tool or AI-assisted development, and 12 (3%) named Codex.
Codex is the example because delegation is what it is best known for, but the method suits any agent that returns a diff. Check the current documentation for settings. For choosing between interactive and delegated tools as a team, see Claude Code versus Codex for engineering teams.
What makes a task safe to delegate?
Three things: it is small, it is clear, and you can check the result without watching the work. A delegated agent works in the background, often in a sandboxed copy of your code, and hands back a diff. You are not there to catch a wrong turn, so the task itself has to make wrong turns obvious.
| Delegate | Keep interactive or do yourself |
|---|---|
| A bug with a reproduction | A bug nobody can reproduce yet |
| A test to add for existing behaviour | Deciding what the behaviour should be |
| A dependency upgrade with a green suite | An upgrade with breaking changes and no tests |
| A mechanical rename or migration across files | A refactor that needs design judgement |
| Documentation for code that exists | Anything in authentication, payments, secrets or data migrations |
If you cannot write down how you will know it worked, the task is not ready to delegate.
How do you write a spec an agent can execute?
Like a good ticket for a capable contractor who has never seen your codebase and cannot ask you questions. A spec that works has five parts:
Goal: Reject orders with a negative quantity at the API.
Context: Validation lives in src/orders/validate.ts; errors use ApiError.
Done when: tests/orders/validate.test.ts "rejects negative quantity" passes,
and the full suite stays green.
Constraints: No new dependencies. Do not edit existing tests.
Out of scope: The admin import path.
The goal says what, not how. The context points to the right files so the agent does not invent a parallel structure. "Done when" is checkable by a machine. Constraints and out-of-scope lines stop the helpful extra change you did not ask for, which is the most common reason a delegated diff gets rejected. When a diff comes back wrong, check the spec first: most bad results trace back to something the spec left open.
Why is a failing test the definition of done?
Because it is the one instruction the agent cannot misread and you cannot argue with afterwards. Before you delegate a fix, reproduce the bug as a test that fails today. Then the agent's job is narrow: make that test pass without changing it, and keep the rest of the suite green.
Two rules keep this honest. First, the assertion belongs to you: tell the agent it must not edit the test, and check the diff to confirm it did not. Second, a passing test proves only what the test checks. If the test is weak, the fix can be wrong and green. Spend your effort on the test and less on the prompt.
Where a unit test is hard to write, a reproduction script that exits with an error works too.
How should you set sandbox and approval settings?
Codex, like most delegation agents, lets you decide what it may do without asking: which files it can write, whether it may run commands, and whether it may reach the network. Start strict and loosen on evidence.
- Writes limited to the workspace. The agent edits its copy of the repository and nothing else.
- Commands it may run unasked: the test suite, the linter, the build. Everything else asks.
- Network off by default, turned on for a task that genuinely needs to fetch a dependency.
- No production credentials anywhere the agent runs. A sandbox is only as safe as the secrets you leave out of it.
Loosen a setting only when you approve the same safe action every time, never to make a task finish faster.
What goes in AGENTS.md?
AGENTS.md is a file of project instructions in your repository that Codex reads before it works. For delegated tasks it matters even more than for interactive ones, because the agent cannot ask you. Keep it to what a newcomer would need on day one: how to install, how to run the tests (the whole suite and one file), the conventions, the folders never to touch, and what to report when finished. Short and specific beats long and thorough. If you also use another agent, keep one source of instructions and point the other file at it.
How do you run several tasks in parallel without chaos?
Delegation pays off when several small tasks run at once and you review the results in a batch. The discipline:
- One task, one branch, one diff. Never let two tasks touch the same files at the same time.
- Keep the queue to what you can review. Five diffs you read properly beat fifteen you skim.
- Review against the spec, not the summary. Does the diff do what "done when" said, and nothing else?
- Use a checklist: tests untouched, suite green, no new dependency, no unrelated change, errors handled rather than swallowed, and a change you could explain in your own words.
- Reject cheaply. A wrong diff is feedback on your spec. Fix the spec and delegate again rather than patching the agent's work by hand.
Reviewing many diffs a day is where control is kept or lost.
What should you never delegate?
- The decision about what to build. Specs are yours.
- The tests that define correctness, or at least their assertions.
- Changes to authentication, payments, permissions or secrets that you cannot verify line by line.
- Irreversible operations: data migrations on real data, deletions, deployments to production.
- The merge. A person who read the diff presses the button.
Where should you start this week?
- Write a one-page
AGENTS.md: install, test commands, conventions, no-go folders. - Pick three small bugs. For each, write a failing test and a five-part spec.
- Delegate all three with writes limited to the workspace and the network off.
- Review each diff against its spec with the checklist above, and note why you rejected any.
- Next week, rewrite the spec template using those rejection reasons, and delegate five.
Square 1's Codex from Zero is an on-demand course, recorded by an instructor and graded by Nova, the AI tutor: seven modules and about 9.5 hours covering sandboxes and approvals, specs an agent can execute, reproduce-then-fix, delegating in the cloud, reviewing many diffs a day, and security and quality passes, ending in a release taken from spec to production. Its projects are a reviewed feature shipped through Codex, ten issues closed on an open-source repository with a failing test first for each, and a release. It asks only that you write software in any language. AI Coding Workflows compares Cursor, Claude Code and Codex on the same feature, and the twelve-week Codex for Software Engineering Bootcamp is live on Zoom with one instructor, about 15 hours a week. All are on a waitlist today. Do software engineering jobs expect AI coding tools? sets out what employers ask for alongside the tools.
