The demonstration goes well. The agent finds the order, checks its status and drafts an answer the support team would happily send. Then someone asks: “What happens if the connection drops just after it changes the order?” For a moment, the discussion stops being about the model.
That question marks a more useful boundary than the end of a proof of concept. Production will bring incomplete requests, unavailable tools and people asking for things nobody anticipated. A CTO needs to know how the service will behave under those conditions and who can take over. The project is ready to move forward when the team can explain that part of the experience as clearly as it can present the successful demonstration.
Be precise about the job you are handing over
“Handle customer enquiries” leaves too much unspecified. Define accepted inputs, available sources, permitted actions and completion conditions. A more useful scope might be preparing an answer about an order using only records that the requesting user is allowed to access. Add explicit exclusions. The agent may provide information without having authority to change prices, amend deliveries or promise compensation. Reflect those boundaries in its tools and permissions as well as its instructions.
Define when it must stop or decline. An unknown identifier, incompatible records or an out-of-scope request requires a controlled response. A useful handover is part of successful service, rather than an embarrassing outcome to hide from the evaluation.
The awkward cases belong in the test set
Collect representative operational examples, with appropriate protection or anonymisation. Include ordinary cases and situations requiring particular care: incomplete data, contradictory documents, unauthorised users and temporarily unavailable tools. For each case, specify the expected result and forbidden actions. Check both the final answer and any changes made to external systems. Courteous wording cannot compensate for an incorrect update.
Anthropic's discussion of agent evaluations distinguishes tasks, trials and grading criteria. A practical implication is to repeat variable cases rather than treating one successful attempt as sufficient evidence. Keep some cases separate from the examples used during development. Otherwise, repeated adjustments can improve performance on familiar tests while leaving new requests poorly handled. Review failures by severity instead of combining every error into a reassuring average.
A proposed action still needs authority
The model can propose an action; the surrounding system must decide whether that action is authorised. This boundary becomes essential when tools can change data or send messages. Use service identities with limited permissions. Validate parameters against business rules and check access to the requested resource. Avoid general-purpose tools that accept arbitrary instructions or operate across unrestricted records.
Treat retrieved documents and messages as data, not as system instructions. A document telling the agent to ignore its policy or transmit information elsewhere must not increase the agent's authority. Human approval should have a precise meaning. Approving a draft does not authorise every later variation of it. If the recipient, amount or commercial conditions change, the system should request a fresh decision where required.
The connection drops. What actually happened?
Suppose the agent sends an update to an external system and loses its connection before receiving confirmation. The agent sees an error; the destination may already have made the change. Pressing “retry” without finding out could repeat an action the customer requested only once.
To recover from that uncertainty, the process needs a stable operation identifier and recognisable states: received, validated, awaiting approval, executed and closed. Keeping a long conversation is not enough. Whoever picks up the case must be able to distinguish attempted actions from confirmed results, then check the destination before repeating an uncertain operation. Where no reliable check exists, an operator needs to handle that uncertainty explicitly.
Approvals can become stale in much the same way. Yesterday's approved proposal may have a different amount or recipient today. Tying approval to the exact proposal, with clear expiry conditions, prevents recovery from executing an old decision against a changed case. This is the kind of detail that rarely makes a demo impressive, yet determines whether interrupted work can safely resume.
Let the service earn greater autonomy
Consider a hypothetical support agent that drafts answers. Its first stage uses historical cases. In the second stage, it proposes answers for live requests while a person reviews every message before sending. A later stage may permit automatic replies for narrow, verifiable categories. Other requests remain supervised. Progress between stages should depend on evidence, rather than an arbitrary commercial deadline.
Keep an operational switch that stops new actions and a clear queue for manual continuation. Rolling back the software does not recall messages already sent or reverse changes made elsewhere. Recovery planning must cover external effects as well as the deployed version. Begin with a limited population so that problems remain manageable. Record why each expansion is justified and what evidence would trigger a return to closer supervision.
Follow the customer through to closure
A technically successful execution can still leave a customer unanswered. Monitoring should follow the business journey from arrival to confirmed closure. Record the case identifier, instruction version, model, tool calls, validation results and outcome. Avoid retaining unnecessary sensitive information. Specify who can inspect traces and how long they remain available.
Combine quality, latency, error rates, human review and cost per completed case. Also watch the age of unresolved cases. A favourable average response time can conceal a small group of requests that have been abandoned for days. For each alert, identify an owner and an action. A dashboard that displays a growing backlog without prompting anyone to intervene is visibility without operational control.
Before launch, someone must know what to do when the model provider is unavailable, a tool returns incomplete records or spending rises unexpectedly. Write short procedures and rehearse them. Define who can stop the service, which cases require escalation and how users are informed. Set spending and execution-duration limits to prevent expensive loops.
Human review needs capacity too. If a pilot sends half of its cases to a reviewer, estimate that workload explicitly. The agent has not increased overall capacity if it transfers more effort than it removes.
Before widening access, bring together the people accountable for business outcomes and technical operations. Work through concrete cases: what the agent may do, how it performed, which permissions it holds and how an interruption would be handled. Record known limitations and the signals that would require closer supervision. During an incident, those decisions are more useful than a general statement that testing went well.
After launch, real failures become candidates for the evaluation set, and changes to models, tools or policies need another check against it. If you bring a proof of concept to DigitalCube, include the case that still worries you alongside the successful demonstration. That question about the dropped connection deserves an answer before a customer has to ask it.

