Signal / 01
Agentic AI watch — Wrong action  •  Looping  •  Lost context  •  Relationship damage  •  Silent drift
Field guide · AI governance

Your AI Agent Is Failing Quietly

The most dangerous contact-center failures do not crash. They complete the wrong task, lose the customer, and report success—unless you instrument the system to see what the dashboard misses.

A contact-center agent wearing a headset faces a customer service screen
The failure that matters may leave no red light—only a confident voice and the wrong result. Photo: 112 Uttar Pradesh / Pexels.

The green checkmark problem

Imagine a customer asks an AI agent to move a payment date. The agent sounds certain, calls a tool, and ends the conversation with a polished confirmation. The workflow returns 200 OK. The contact is marked resolved. But the date did not move; the tool updated a different field. Every machine in the chain says success. The customer discovers the truth when a late fee appears.

This is the failure that matters now. Contact-center AI is moving from answering questions to taking actions, and action changes the risk. A fluent wrong answer is irritating. A fluent wrong action can alter an account, issue money, cancel service or create a compliance problem. The most dangerous incidents rarely arrive as outages. They look like completed work.

That is why autonomy without governance is liability with a confident voice. Governance is not a policy document beside the system. It is the set of controls inside the flow: what the agent may do, what evidence counts as success, when a human must enter and what gets recorded so the decision can be reconstructed later.

It completes the wrong action confidently

An agent can understand the sentence and still misunderstand the job. It can select the wrong account, confuse a refund with a credit, or call the right tool with the wrong parameters. The conversational layer remains smooth because fluency and transaction correctness are different systems.

Log: the customer’s stated intent, the normalized intent chosen by the agent, every tool name and parameter, the authorization state, the policy version, the tool response, the before-and-after business state, and the final words shown to the customer. Give each interaction a trace ID that survives every hop. Redact sensitive fields, but do not reduce the record to a transcript and a success flag.

Alert: an agent says “refund issued” without a matching ledger event; the requested amount differs from the executed amount; a write occurs on an entity not verified in the session; or the state has not changed after a claimed success. High-consequence actions need deterministic controls outside the model—limits, allowlists, identity checks and, where necessary, human approval. The model can propose. Code decides what is permitted.

It loops without escalating

The loop is the oldest chatbot failure in a smarter suit. The customer rephrases. The agent apologizes. The customer rephrases again. Each turn looks locally reasonable, so no individual message trips an alarm. The failure exists in the sequence.

Log: turn count, repeated-intent similarity, repeated tool calls, unchanged state, fallback phrases, sentiment trajectory and explicit requests for a person. Track progress as an event: did the interaction acquire new information, change state or narrow the problem? A polite exchange with no progress is still a loop.

Alert: when the same intent is detected repeatedly without state change, when an identical tool fails twice, when sentiment falls across consecutive turns, or when a customer asks for a human and the next action is anything but escalation. Put a hard ceiling on retries. Escalation should be a reachable state in the workflow, not a suggestion left to a probabilistic model.

A human customer-support agent concentrates during a headset call
A human escalation is only useful when the person receives the context needed to act. Photo: Yan Krukau / Pexels.

It escalates and drops the story

A transfer can be operationally successful and experientially broken. The queue receives the contact on time, but the human receives no useful history. The customer starts again: identity, problem, attempted fix, emotional temperature. The company counts a handoff. The customer counts abandonment.

Log: the transfer reason, detected intent, verified identity, entities discussed, actions attempted, tool results, promises made, sentiment, channel history and the summary delivered to the human. Then log whether the agent desktop actually rendered that context and whether the employee reopened the same questions. Context delivery needs its own acknowledgment, just like a financial message.

Alert: on empty or unusually short summaries, missing required fields, transfers without a reason code, repeated authentication immediately after handoff, and a sharp rise in handle time following an AI transfer. UJET’s live Virtual Agent product emphasizes seamless human handoff with journey context. That claim points to the right operational standard: the transfer is not complete when the call moves; it is complete when understanding moves with it.

The transaction succeeds and the relationship loses

Suppose the agent correctly denies a fee reversal. The policy was followed, the CRM was updated and the case was closed. The customer, calling after a family emergency, leaves convinced that the company did not listen. The system has succeeded in the narrowest possible sense.

Contact centers inherit a measurement habit from software operations: if the function returned the expected value, the function worked. Relationships are not functions. A correct action can still be badly timed, needlessly rigid or emotionally tone-deaf.

Log: the sentiment trajectory rather than a single final score; interruption and repetition; time spent contesting a decision; escalation offers; complaint language; cancellation or churn signals; and post-contact behavior such as a repeat contact within 24 hours. Segment these outcomes by intent and customer context.

Alert: when transaction success rises while customer effort, repeat contact, complaints or negative sentiment worsen. That crossed signal is the tell. A dashboard that only reports containment will reward the machine for keeping an unhappy person away from a human. The metric should be durable resolution, not merely a conversation that ended.

The crossed signal

Transaction success↑ rising
Customer health↓ falling

A dependency changes and quality quietly decays

Agentic systems sit on moving ground. A CRM field is renamed. A knowledge article changes. An authentication rule tightens. A model version shifts. The agent keeps speaking, so the failure does not resemble downtime. It resembles a small decline in successful writes, a few more retries and a subtle increase in repeat contacts.

Log: versions for the model, prompt, policy, knowledge source, tool schema and downstream APIs on every trace. Record retrieval results, tool latency, error classes, token and cost changes, and outcome metrics by version. Maintain synthetic journeys—stable test customers with known answers—that run continuously across the highest-risk flows.

Alert: on change points, not only fixed thresholds: a drop in tool success after a schema release, a shift in intent distribution after a prompt update, rising latency, retrieval returning older documents or disagreement between synthetic expected states and actual states. Every dependency change should have an owner, a canary cohort and a rollback path. If you cannot say what changed between Tuesday and Wednesday, you do not have an observable agent; you have a séance.

A contact-center professional listens carefully while working at a computer
The hardest cases bend with context—the place human judgment matters most. Photo: Antoni Shkraba / Pexels.

The autonomy test

The useful question is not whether the agent is intelligent enough. It is how much authority this particular flow deserves. Ask two things. First: who suffers if it fails? Second: is the correct answer fixed, or does it bend with context?

A store-hours lookup has a fixed answer and a low cost of failure. Automate it. A small refund may have clear rules but a financial consequence. Automate within coded limits and audit every write. A cancellation request from a distressed long-term customer has an answer that bends with context and a high relationship cost. Assist a human or require approval. Medical, safety, fraud and vulnerable-customer cases belong behind the strongest boundaries.

This produces a simple autonomy map: low consequence plus fixed answer can run; high consequence plus fixed answer runs only behind hard controls; low consequence plus contextual answer needs monitoring and an easy exit; high consequence plus contextual answer needs a person in the decision. The unit of governance is the flow, not the brand promise.

Build the alarm before the agent

UJET Virtual Agent is in production today, connecting to back-end systems, supporting more than 40 languages and routing between virtual and live help using context. UJET announced Agentic Experience Orchestration, or AXO, on March 11, 2026. AXO is scheduled for general availability at the end of September 2026; it should be understood here as announced, not yet deployable. Its framing—variable AI control, humans in the loop, persistent context and continual assessment—matches the larger lesson.

But no platform absolves an operator of design. Before granting an agent a new tool, define the prohibited action, the evidence of success, the retry limit, the escalation payload and the owner of the alert. Replay real interactions. Run the same task repeatedly. Break dependencies on purpose. Test the handoff from the human’s screen, not the AI’s transcript.

The honest promise of agentic AI is not that it never fails. Complex systems fail, and systems that converse can hide the failure unusually well. The promise worth making is that failure will be bounded, visible and recoverable. The best-run contact center will not be the one with the most autonomous agent. It will be the one that knows, quickly and precisely, when autonomy has stopped helping.