Skip to content
Back to the blog

AI agents: what must change between a pilot and production?

How to evaluate an AI agent before it joins a live business workflow: task boundaries, tool permissions, failure handling and operating ownership.

AI in practiceAbout 6 min read
A charcoal route crosses a dashed boundary at a circular checkpoint, with an amber branch curving back to a second circle
Editorial illustration

Define the job the business is relying on

An AI pilot can look complete when it answers a few prepared questions. Production introduces work the demonstration did not settle: incomplete records, unavailable systems, conflicting instructions and people who expect a request to be handled. The launch decision should describe the responsibility the organization is ready to give the agent.

Start with a task statement that names the input, allowed sources, permitted actions and completion condition. “Help the support team” is too broad. “Read an assigned request, retrieve the relevant order and prepare a response for review” can be tested. Decide who owns requests the agent cannot finish, and how those requests appear in the team’s normal queue.

For buyers evaluating agentic AI solutions in Israel, language is part of this definition. If work arrives in Hebrew and English, include both, along with mixed-language messages, product identifiers and local date conventions. A fluent answer does not show that the agent selected the correct customer record.

Turn pilot observations into launch evidence

Turn pilot observations into launch evidence
A pilot may showProduction needs evidence ofA useful check
A plausible answerAn answer supported by the permitted source.Remove the relevant source and inspect the response.
A successful tool callThe correct action on the correct record.Check the destination system after execution.
One completed scenarioConsistent handling of representative cases.Repeat cases and review failures by category.
A request for human helpA handover someone can act on.Open the assigned queue and inspect its context.

Separate access to information from authority to act

List the actions separately: retrieve an order, draft an internal note, update a status, send a message. Each can require a different permission or review step. Keep authorization checks in the application that executes the action. Instructions telling a model to behave carefully do not replace a system checking whether that action is allowed.

Anthropic’s guidance on building agents emphasizes clear tool definitions and testing their use. For a buyer, the practical question is whether each tool has a narrow, understandable purpose. Ask your supplier to demonstrate what happens when a required identifier is missing, a record belongs to another account or the requested action is outside the agreed scope.

Also decide when the agent stops. A failed lookup should not cause endless attempts or an improvised answer. Define limits on repeated calls, an escalation path and the information retained for the person who takes over. Choose those limits around the task rather than adopting a generic autonomy setting.

Build a test set around the consequences of being wrong

Collect representative cases from the process, with appropriate permission to use the data. Include ordinary requests and cases where the right result is to ask for clarification or stop. Separate a wrong answer from a wrong action: sending a confident explanation and changing the wrong record have different consequences and should not disappear inside one average score.

Anthropic’s evaluation guidance distinguishes the record of an agent’s steps from the resulting state. Inspect both. A response saying “updated” should be checked against the actual record. Repeat selected cases because model behavior can vary between attempts, and review why a case passed rather than relying only on its score.

Keep some cases outside day-to-day tuning, then rerun them before a release. Document which model, prompts, tools and source version were tested. This makes later changes comparable and helps identify whether a regression came from the agent or from information and systems around it.

A hypothetical rollout: order-support assistance

Imagine a distributor whose team receives questions about open orders. In the first release, the agent reads an assigned request, finds the order and prepares a draft with links to its evidence. A team member approves the response. The agent cannot change delivery dates or issue credits.

The test set includes two customers with similar names, an order missing from the source, an outdated delivery estimate and a request that contains instructions to ignore the normal process. The team checks correct record selection, supported answers and useful escalation. Only after that evidence is reviewed should it consider granting a narrowly defined additional action.

Start live use with an audience the operating team can support. Observe actual corrections and unresolved cases, then add them to the test set. Expanding the audience and expanding the agent’s permissions are separate decisions; each deserves its own review.

Ask for the operating plan before launch

The agent will depend on changing data, tools and business rules. Assign responsibility for those changes alongside the initial implementation.

  • Who reviews incorrect answers, unauthorized attempts and failed actions?
  • Can an operator pause execution while requests remain available for manual handling?
  • What evidence is retained, who may view it and how is unnecessary customer data avoided?
  • Who approves changes to prompts, models, tool access and source material?
  • How are task completion, human corrections, response time and operating cost reviewed together?

Questions to settle with your development partner

Does production mean full autonomy? No. A system that prepares reliable work for approval can be a complete production release. Choose the authority around the task and the organization’s capacity to review it.

Can a stronger model replace evaluation? It still needs to be tested in your workflow. Model capability does not establish that your source data, permissions or integration behavior are correct. Ask for a repeatable evaluation process and a demonstration of failure recovery, alongside the successful case.

Sources and further reading

Define what your agent can be trusted to do.

Bring the workflow, available systems and examples of difficult requests. We can discuss a focused scope and the evidence needed before live use.

Discuss an AI agent project
Nexo