One-minute overview
- The dividing line for an agent is not "can it work once" but "can it keep running reliably, reversibly and auditably": demo performance is not production reliability.
- Screen tasks first with three rulers — determinism, reversibility, feedback loop. Start with low-risk work that has a clear right/wrong answer and can be verified automatically; don't start by sending emails, changing databases or making payments.
- Don't evaluate on a single success rate: you need the trio of task completion rate + trace evaluation + a reproducible regression set. Failure samples expose problems better than a pass rate does.
- Human-in-the-loop is not "adding a confirm button" — it is approval gates, least privilege and rollback working together, priced by risk.
- Build governance from day one: per-task cost, error amplification and audit logs are all mandatory. Quote vendor positions verbatim, verify every one, and don't accept unverifiable "high success rate" slogans.
The analogy in one line
Moving an agent from demo into production is like turning a "sample that ran fine once" into a "factory assembly line": the sample only has to look good, while the line is judged on yield, downtime, traceability and scrap rate — none of those four metrics is proven by "it worked once."
1. Introduction: agents move from concept to production
The AI Agent in this article means a system capable of planning, calling tools and multi-step execution: it can break a goal down, decide for itself which tool to call next, and keep chaining actions until the task is finished. It is not a turn-by-turn chatbot, and it is not a fixed pipeline hard-coded into software (that is a workflow).
The concept itself is not new, but "from concept to production" is exactly where most teams are stuck today. The demo stage asks "can it work at all?"; the production stage asks four other things:
- Reliability: if you run the same task 100 times, what does the distribution of results look like? Will it be right once and wrong the next time?
- Observability: what did it do at each step, which tool did it call, why that choice — can you trace it back?
- Rollback: if it gets something wrong, can you undo it at low cost?
- Governance: do permission boundaries, cost caps and audit records exist from day one?
Vendor definitions happen to draw the boundary at that same action level of "tool use." ⚑Vendor position — Anthropic defines it in "Building effective agents" (2024-12-19):
"Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks." (Source: Anthropic engineering article "Building effective agents", 2024-12-19)
Note the two keywords in that sentence: dynamically direct and maintaining control. The first is agency; the second is controllability. Achieving only the first is a demo; achieving both is production.
Fact-check note: the article's page header now carries a banner (translated from the Chinese text) stating: "Since December 2024 the toolchain described here has changed considerably; for current practice see Claude Managed Agents and its accompanying documentation." Quoting its definition and the workflow/agent distinction remains valid, but framework-selection advice must follow the vendor's latest documentation.

2. Capability boundaries: which tasks to hand to an agent first
Start with a counter-intuitive answer: it is not the most "intelligent"-sounding tasks that are ready for an agent first, but the most "verifiable" ones.
Use three rulers to screen candidate tasks:
| Ruler | Question to ask | Good signal | Bad signal |
|---|---|---|---|
| Determinism | Does this task have a clear right or wrong? | A standard answer, or an outcome that can be judged automatically | Quality depends on subjective judgement |
| Reversibility | If the agent gets it wrong, can it be undone at low cost? | Read-only, draft, reassignable | Already sent, already charged, database already changed |
| Feedback loop | Is there an automatic mechanism that tells it "right or wrong"? | Tests, validation, human labels can flow back | Errors can only be found manually after the fact |
Applying those three rulers, the practical candidate list looks like this:
| Task | Determinism | Reversibility | Feedback loop | Verdict |
|---|---|---|---|---|
| Ticket classification and routing | High | High (just reassign) | Yes (humans can label right/wrong) | Good first task |
| Code generation + passing unit tests | Medium | High (commit can be reverted) | Yes (tests are the feedback) | Suitable, but tests are mandatory |
| Meeting notes + action-item extraction | Medium | High (humans can proofread) | Yes (proofreading flows back) | Suitable, output goes to human review |
| Support reply drafts | Medium | High (drafts only) | Yes (human gate before sending) | Suitable, but drafts only |
| Automatically sending outbound email | Medium | Low (once sent, it's gone) | Weak | Start with "draft + human sends" |
| Bulk-editing the production database | Low | Low | Weak | Not suitable; at minimum read-only + approval first |
A plain conclusion: "read-only first, drafts first, tests first" is the shortest path to putting agents into production. The three categories of actions that face the outside world, touch money, or touch production data all come later.

3. Reliability and evaluation: how to measure "good enough"
Production reliability cannot rest on gut feeling; it rests on a trio, and all three are required:
1. Task completion rate — but it must be broken out, not aggregated. Reporting one overall success rate means nothing. Break it down by task type, difficulty tier and tool combination, and state the test set, the number of samples and the evaluation method. A "90% success rate" that doesn't say which test set it was measured on, how many runs it took, or how right and wrong were judged is as good as unreported.
2. Trace evaluation — look at the process, not just the outcome. Many tasks get "the right answer" while "the process was a guess." Evaluate whether each step's tool choice was reasonable, whether it took detours, whether it made out-of-scope calls. Trace evaluation catches the hidden failures where "the result was right but the path was wrong" — and in production those failures eventually blow up.
3. Regression test set — to prevent "fix one thing, break a dozen." Distil historical failure samples into a regression set, and run it every time you change a prompt, swap a model or adjust a tool. The regression set must be reproducible: same input, same environment, same scoring rules, same result.
Practical advice: public benchmarks (SWE-bench, GAIA, tau-bench and the like) are for cross-comparison only; they do not replace your own business regression set. They don't cover your real tools, your real data or your real failure modes. What is genuinely valuable is the pitfalls your own team has stepped in — freeze those into the regression set.

4. Human-machine collaboration: design trade-offs of human-in-the-loop
The goal of HITL is not "have a person watch every action" (that throws away the value of automation) but to spend human attention where the risk is highest and the judgement hardest. Three design elements:
1. Approval gates — tiered by risk, not one-size-fits-all. Low-risk actions (reading data, generating drafts) run automatically; medium-risk actions (changing configuration, bulk operations) run after approval; high-risk actions (outbound sending, funds, writes to production data) are performed by humans. The risk tiers belong in configuration, not in a developer's ad-hoc judgement.
2. Least privilege — give the agent only what the task requires. An agent's toolset should be explicitly defined: databases it can read should not grant write permission, and an account that can send drafts should not have broadcast permissions. A permission violation is itself a failure signal and should raise an alert.
3. Rollback — leave a way back at every step. Write operations need a corresponding undo path; outbound actions must be retractable or turned into drafts first. Rollback is not a remedy; it is a precondition for shipping.
The balance point: automation benefit = human cost saved; HITL cost = approval latency + human backstop. When the human-intervention rate for a task type is stable over time and the error rate acceptable, consider loosening approvals; otherwise tighten. Approval thresholds should follow the data, not be frozen at launch.

5. Cost and governance: token cost, security boundaries and audit
1. Per-task cost accounting — count the whole chain, not a single call. Per-task cost = (input tokens + output tokens + extra tokens from tool calls/retrieval) × unit price + cost of the number of tool calls + human backstop cost. Multi-step agents quietly burn money on "calling repeatedly, re-reading context"; watching only the price of one model call badly underestimates the real cost.
2. Error amplification risk — cascading failure in multi-step tasks. In a single-step task, one wrong step is just one wrong step; in a multi-step task one wrong step can cascade (a bad intermediate result is fed into later steps). Mitigations: put checkpoints at key steps, set a task-level budget cap (tokens or steps), and trip a circuit breaker to hand over to a human when it is exceeded. The circuit breaker should be designed as a normal mechanism, not as an emergency fallback.
3. Audit logs — record from day one. Each run should record at minimum: the input (after de-identification), the model output and tool calls at every step, the permissions and approvals triggered, and rollbacks and human interventions. Audit logs are not just for compliance; they are the raw material for post-mortems, for training the regression set and for costing.
Governance checklist:
- Every agent task has an explicit permission list and cost cap
- Write operations and outbound operations have approval gates and rollback paths
- Audit logs cover the whole chain: input, tool calls, approvals, rollback
- Failure samples feed the regression set automatically, and evaluation is reproducible
- Per-task cost is reviewed monthly, with alerts on abnormal increases

6. Vendor positions and roadmap
Newsroom note: every vendor statement in this section has been checked against the official original (verification date at the end), and vendor positions are uniformly marked ⚑. Vendor quotes are reproduced verbatim without extrapolation; before publishing, click through the source links once more (official documentation is updated continuously).
⚑ Anthropic makes two key judgements in "Building effective agents" (2024-12-19):
"Workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale." (Source: Anthropic engineering article "Building effective agents", 2024-12-19)
That is: when tasks are clearly defined and the path is fixed, workflows are more predictable and more stable; agents come in when flexibility and model-driven decision-making are needed. ⚑Anthropic also summarises customer practice: the most successful implementations tend to use simple, composable patterns rather than complex frameworks or purpose-built libraries (same source; paraphrase of the gist; the article's header now notes the toolchain has changed, see the fact-check note in section 1).
⚑ OpenAI gives a definition in "A Practical Guide to Building Agents" (official PDF, published 2025, 34 pages):
"Agents are systems that independently accomplish tasks on your behalf." (Source: OpenAI, "A Practical Guide to Building Agents", official PDF)
and two operational characteristics:
"It leverages an LLM to manage workflow execution and make decisions... It has access to various tools to interact with external systems—both to gather context and to take actions—and dynamically selects the appropriate tools depending on the workflow's current state, always operating within clearly defined guardrails." (Same source; the ellipsis marks omitted text from a longer passage, not a rewrite)
The point of that wording: an agent uses an LLM to manage workflow execution and decisions and dynamically selects tools within clearly defined guardrails — which is the same thing as this article's governance stance of "least privilege + approval gates," stated from the vendor side.
⚑ Microsoft is pushing at the toolchain level: AutoGen / Semantic Kernel provide agent programming frameworks, Copilot Studio provides low-code build, test and publish, and within the Azure AI Foundry ecosystem it is advancing agent services for production hosting (paraphrase of Microsoft's position, not a verbatim quote; specific product names and availability are subject to Microsoft's official documentation).
Key data table: vendor positions and production reality, side by side
| Topic | Vendor position (verbatim or paraphrase, with source) | Production reality / adoption judgement |
|---|---|---|
| What an agent is | ⚑Anthropic: "Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks." (2024-12-19) | "Dynamically directing + calling tools" is only the starting point; production also needs observability, rollback and audit |
| When to use an agent | ⚑Anthropic: workflows suit well-defined tasks; agents suit scenarios where flexibility and model-driven decision-making are needed at scale (2024-12-19) | When task boundaries are clear and the path is fixed, use a workflow first; don't adopt an agent for the sake of the word |
| Build complexity | ⚑Anthropic: the most successful implementations use simple, composable patterns (paraphrase of the gist, 2024-12-19) | Start with a single agent plus an explicit flow; add complexity only after the benefit is proven |
| Agent definition (OpenAI) | ⚑OpenAI: "Agents are systems that independently accomplish tasks on your behalf." ("A Practical Guide to Building Agents" PDF, 2025) | "Independently accomplishing tasks" is the goal; production also needs guardrails, human review and rollback |
| OpenAI guardrails | ⚑OpenAI: agents dynamically select tools "within clearly defined guardrails" (same source) | Toolset, call timing and guardrails must be explicit — that is the least-privilege interface layer |
| Microsoft ecosystem | ⚑AutoGen / Semantic Kernel / Copilot Studio / Azure AI Foundry toolchain (paraphrase, not verbatim) | Choose a framework by whether it makes evaluation, tracing and permission control easy — not by how easy the demo is to write |
| "High success rate" claims | Vendor demos commonly cite "high success rates", mostly without disclosing the test set or sampling method (unverified, not accepted) | Go by reproducible evaluation: task completion rate + trace evaluation + your own regression set |

7. Conclusion: practical advice for the newsroom and metrics to track
The priority advice for putting agents to work, in one sentence: pick "read-only / draft / verifiable" tasks first, evaluate with the trio, design human-machine collaboration by risk tier, and stand up cost and audit on day one.
Metrics to keep tracking:
- Task completion rate: broken out by task type and difficulty tier, with the test set and sampling method stated
- Trace pass rate / key-step pass rate: look at the process, not a single outcome
- Human intervention rate: HITL trigger rate and approval rejection rate, used to tune approval thresholds dynamically
- Full-chain per-task cost: tokens + tool calls + human backstop, reviewed monthly
- Number of rollback / failure events and recovery time (MTTR): measures "how fast we recover when something breaks"
- New failures added to the regression set: measures "did fixing one thing break a dozen"
- Number of error-amplification events: incidents where one error caused multiple downstream failures, used as a signal for tuning the circuit breaker
Sources to keep tracking:
- OpenAI: official blog and documentation ("A Practical Guide to Building Agents" PDF, plus the Agents SDK / Building agents series)
- Anthropic: official engineering articles ("Building effective agents" and subsequent Claude Managed Agents updates)
- Microsoft: official documentation for Azure AI / AutoGen / Semantic Kernel / Copilot Studio
- Public benchmarks: updates to SWE-bench, GAIA, tau-bench and others (for cross-comparison only)
- Vendor safety cards / system cards: capability boundaries and risk statements
Closing line: an agent's productivity does not lie in how smart it is, but in whether it can afford to be wrong, be seen, be taken back and be audited. Those four things are the real roadmap from demo to production.
Fact-check note (review editor)
- Verification date: this review. Anthropic's two verbatim quotes were checked word by word against “Building effective agents” (including "on the other hand"); the "workflow vs agent" sentence matches verbatim; "simple, composable patterns" is a paraphrase of the gist.
- The OpenAI quote was checked word by word against the official “A Practical Guide to Building Agents” PDF (34 pages, published 2025), replacing statements and dates in the first draft that could not be sourced.
- The Microsoft toolchain description is a paraphrased vendor position, not a verbatim quote, and is marked ⚑.
- SWE-bench, GAIA and tau-bench are public industry benchmarks used here for cross-comparison only; no specific scores are cited.

