You build a production app on the Claude API by putting the API behind your own server, giving the model a versioned system prompt and a defined output shape, streaming answers to the user, retrying the errors that are worth retrying, caching the stable parts of your prompts, logging every call with its cost, and running an eval suite before every change. Skipping any one of them is how a good demo becomes a bad product. Employers are hiring for this work: in 315 AI and machine learning job ads we collected on 19 September 2026, 49 (16%) named Anthropic or Claude and 28 (9%) named the OpenAI API.
This piece describes the architecture in general terms. The API changes often, so check the current documentation for model names, prices, limits and exact parameters before you build.
What does the architecture look like?
Three layers, and the middle one is yours:
| Layer | What it does | What it must never do |
|---|---|---|
| Client (browser or app) | Collects input, shows streamed output, lets the user cancel | Hold the API key or call the API directly |
| Your backend | Authenticates the user, applies your own limits, assembles the prompt, runs tools, logs, handles errors | Trust model output or tool output without checking it |
| The Claude API | Generates the response, requests tool calls, returns usage figures | Receive data it does not need |
A thin backend that only forwards requests gives you none of the protections below.
Where should the API key live?
On the server, and nowhere else. Keep it in an environment variable or a secrets manager, never in the client bundle, a mobile app or the repository. Use separate keys per environment, and know how to rotate one quickly. Put your own per-user limits in front of the API, so one enthusiastic user or one runaway loop cannot exhaust the capacity everyone else depends on.
How do you design the prompt and the output?
Treat the system prompt as code: keep it in version control, log its version with every request, and change it only with an eval result in hand. Rules, refusals and output format go in the system prompt; the user's material goes in the messages.
When your code needs to act on the answer, do not parse free text. Ask for structured output, either through a schema or through tool use, and validate it in your code. With tool use, the model does not run anything itself: it asks your code to call a tool, your backend checks the request and the user's permissions, runs it, and returns the result for the model to use. That loop, and how to write tool descriptions the model understands, is covered in our explainer on tool use and function calling. Anything irreversible, such as sending an email, issuing a refund or deleting a record, should wait for a person to approve it.
How do you handle streaming, errors and rate limits?
Streaming. For anything a person watches, stream the response through your backend so text appears as it is generated, and support cancellation so a user who has seen enough stops paying for the rest.
Errors. Sort them into two groups. Rate-limit responses, overloaded or temporary server errors, timeouts and dropped connections are worth retrying. Invalid requests and authentication failures are not. Retry with exponential backoff and random jitter, cap the number of attempts, and honour any retry guidance the response gives you. Set timeouts on every call so a slow request fails cleanly instead of hanging a user's session.
Rate limits. Your account has limits on requests and tokens. Design for them rather than discovering them in production: queue work that does not need an instant answer, smooth out bursts, and use the API's batch processing for large jobs that can wait. If a tool call has side effects, make it idempotent, so a retry does not charge a card or send an email twice.
How do you keep the cost under control?
Cost grows with tokens, so the levers are the tokens you send and receive:
- Prompt caching. If many requests share a long stable prefix, such as the system prompt, tool definitions or a reference document, put that material first and mark it for caching, so repeated requests can reuse it at lower cost and latency. Check the documentation for how caching is priced and how long a cache lasts.
- The right model for the job. Use the smallest model that passes your eval, and send only the hard cases to a larger one.
- Less in, less out. Trim long conversation histories, retrieve only the passages a question needs, and cap output length.
- Measurement. The API reports token usage on each response. Record it per request, per feature and per user, and set a budget alert before the invoice sets one for you.
What should you log, and what should you keep out?
Log enough to explain any answer after the fact: a request identifier, the prompt version, the model, token counts, latency, every tool call and its result, the outcome and any error.
Keep personal information out of prompts and logs unless the feature genuinely needs it. Redact identifiers where you can, set a retention period for logs, restrict who can read them, and read the provider's data terms against your own obligations, such as the Privacy Act if you handle Australians' information. Our piece on data privacy in AI systems covers the wider picture.
How do you know it is ready to ship?
With evals and a safety pass. Build a golden set of real inputs with expected results, score the application on it, and run that suite automatically whenever the prompt, the tools or the model change; a drop in the score blocks the release. Our guide to evaluating LLM outputs explains how.
For safety, assume that anything the model reads, including user messages, uploaded documents and tool results, may contain instructions written by someone else. That is prompt injection, and the defence is architectural: give tools the least privilege they need, check every tool call in code, keep a person in the loop for consequential actions, and never let model output reach a database query or a shell without validation.
Where should you start this week?
- Move any API call out of client code and behind a small backend endpoint with the key in a secret.
- Put your system prompt in its own file under version control and log its version with every request.
- Wrap the call with timeouts, capped retries with backoff, and a clear error for the user.
- Log token usage per request for a week, then decide what to cache and what to trim.
- Write a first golden set of real inputs and run it before your next prompt change.
Where does Square 1 teach this?
Building with the Claude API is an on-demand course recorded by an instructor and graded by Nova, the AI tutor: about twelve hours across eight modules, for developers in Python or TypeScript. The graded work runs from a first request with error handling, through an assistant with three tools, a streaming chat interface, a cost comparison with caching on and a document question-answering feature with citations, to a deployed application with its eval suite and a cost report. MCP and Tool-Using Apps goes further into exposing and consuming tools the standard way. The Gen AI Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, with the Claude and OpenAI APIs in its first block and a production LLM product with monitoring and a cost model as its capstone. All three are taking a waitlist today. The full-stack engineer role page shows where this work sits in a product team.
