Companies
23/09/2026

Meta’s Human Backup Reveals the Limits of AI Phone Agents




Meta's decision to test human contractors behind its new Muse personal AI agent exposes one of the most difficult problems facing the next generation of artificial intelligence: making autonomous systems work reliably in the unpredictable real world. Muse is designed to move beyond answering questions and actually perform tasks for users, including making phone calls to businesses. Yet the need to temporarily involve human workers suggests that the technology still struggles when conversations become complicated, resistant or socially unpredictable.
 
The development is particularly significant because Meta has positioned Muse as a personal AI agent rather than another conventional chatbot. The system is intended to browse websites, complete forms, book appointments, make purchases, send emails and handle other multi-step activities on behalf of users. Meta says Muse operates inside a dedicated secure virtual machine and that users retain control over connected services and sensitive actions.
 
The telephone feature takes that ambition into a much harder environment. Websites provide structured information that software can often interpret systematically. Human conversations are different. People interrupt, misunderstand, refuse requests, ask unexpected questions and sometimes react negatively when they discover they are speaking to an automated system. The reported human concierge experiment appears to have emerged from precisely this gap between what an AI agent can theoretically do and what it can reliably accomplish in everyday interactions.
 
Phone Calls Expose AI’s Reliability Gap
 
Muse's phone-calling capability is designed to allow users to delegate ordinary errands to the system. A user can ask the agent to contact a business, check product availability, negotiate a bill, obtain a quotation or arrange an appointment. That changes the nature of the technology because the AI is no longer simply generating information for a person to evaluate. It is representing that person in an interaction with another human being.
 
That distinction creates a much higher reliability requirement. If an AI gives an imperfect answer to a general question, the user can usually recognise the problem and ask again. A failed phone call can have a more direct consequence. A booking may not be completed, an incorrect request may be communicated to a business or a negotiation may produce an unintended result.
 
Reports that some businesses were hanging up when they recognised that Muse was an AI caller underline the difficulty. The challenge is not necessarily that the system cannot understand language. It is that the other participant may refuse to engage with an automated caller altogether. Similar resistance could become a recurring problem for any AI agent that attempts to replace human interactions rather than merely assist them.
 
This helps explain why human intervention can produce dramatically higher completion rates during testing. According to internal information reported about the experiment, calls handled by human contractors could achieve success rates as high as 95 percent to 98 percent, compared with lower rates for fully automated calls. The comparison is revealing because it measures not only the intelligence of the software but also the complexity of the environment in which the software operates.
 
The Human Concierge Creates a Trust Problem
 
The most difficult issue is not necessarily the use of human labour itself. Human assistance is common in technology services, customer support and other digital businesses. The problem arises when users believe they are delegating a task to an autonomous AI agent but a human worker may actually be handling part of the interaction without their knowledge.
 
That distinction becomes especially important for a personal agent because the system may have access to information about a user's life that ordinary customer service software would never possess. Meta says Muse can connect to email, calendars and other applications, while its secure virtual machine stores data and credentials associated with connected services. The company also says Muse does not share the contents of that virtual machine with its advertising systems.
 
A human contractor entering that process introduces a different category of exposure. Even if the contractor receives only the information necessary to complete a call, the interaction could reveal names, addresses, account details, spending information, medical or insurance information and other sensitive circumstances depending on the task.
 
Internal employee criticism reportedly focused precisely on this concern. One employee described the possibility of accidental information disclosure, while another reported an offensive remark by a human contractor during a call involving a service provider. Meta acknowledged that launching such a test without adequate disclosure was a mistake and temporarily rolled the feature back.
 
The episode demonstrates that privacy cannot be treated simply as a technical problem involving encryption and isolated computing environments. It is also a question of who is allowed to observe information when an AI agent cannot complete a task by itself.
 
Meta’s Security Architecture Faces a Practical Test
 
The human concierge experiment is particularly notable because Meta has made security a central part of Muse's design. The company created a dedicated virtual machine for each user, with separate storage for credentials and controls intended to prevent the agent from accessing passwords and payment information directly. A separate security system is designed to approve actions before the agent reaches the internet. Meta has also said that it plans to introduce a confidential version of the virtual machine in which the user's data and conversations would be encrypted with a key controlled by the user. The company says this architecture is intended to prevent even Meta from accessing information stored in that environment.
 
Those protections address an important class of cybersecurity risks, but they do not automatically solve the problem created when an AI task is transferred to a person outside the secure environment. A virtual machine can isolate software from other systems; it cannot make a human conversation private once information has been deliberately exposed to another person to complete a task. That distinction could become increasingly important as AI agents move from digital environments into the physical economy.
 
The Muse experiment also raises a broader question about how the AI industry measures automation. A service may appear autonomous from the user's perspective while relying on human workers for difficult cases behind the scenes. Such arrangements can be useful during product development because they allow companies to identify failure points and understand how humans solve problems that machines cannot yet handle.
 
There is precedent within the technology industry. Meta itself previously experimented with a digital assistant known as M, which reportedly depended extensively on human workers to complete tasks. The current Muse experiment therefore illustrates a recurring pattern: the path towards autonomous AI can initially require more human involvement than the final product suggests.
 
This does not make the technology meaningless. Human assistance during development can help train systems, improve workflows and identify situations that require better safeguards. But the distinction between testing and genuine autonomy becomes important when companies make claims about what an AI agent can independently accomplish.
 
For investors and users, the practical question is therefore not simply whether a company has created an AI capable of performing a task under ideal conditions. It is how often the system can complete that task without human intervention, how failures are handled and whether users are informed when another person becomes involved.
 
The Agent Economy Depends on Cooperation
 
Muse's difficulties also point towards a problem that may affect the entire AI agent industry. An autonomous agent can be technically capable of calling a business, making a purchase or completing a form, but the external service does not have to cooperate. Amazon's recent decision to block Muse from accessing its platform illustrates this wider problem. Amazon said the agent violated its conditions of use, while the dispute raised questions about whether third-party AI systems should be permitted to browse websites and conduct transactions on behalf of users.
 
This means the future of AI agents will depend partly on standards established by the businesses they interact with. If companies refuse automated access, require explicit identification or restrict machine-driven transactions, an agent may need alternative methods to complete tasks. Phone calls present the same problem in a different form. A business can simply refuse to speak to an automated caller. The technical sophistication of the AI cannot force the other party to participate.
 
The issue matters because Muse has attracted significant early attention. The app has surpassed 2.5 million downloads according to data cited in the original report, while other market intelligence has also shown rapid adoption following its launch. Its popularity suggests that consumers are interested in AI that performs tasks rather than simply generates text.
 
That creates pressure on Meta to make the system dependable at scale. A conversational assistant can tolerate a certain amount of imperfection because the user remains directly involved. A personal agent assumes responsibility for actions, which raises expectations around accuracy, transparency and accountability.
 
The human concierge experiment therefore provides an unusually clear indication of where the technology stands. AI agents have progressed considerably in their ability to navigate digital environments, but real-world autonomy remains more difficult because success depends on other people, unpredictable conversations and external systems.
 
For Meta, the challenge is consequently not just to make Muse more intelligent. It is to establish when the system can safely act alone, when it needs assistance, and how clearly that transition is communicated to the user. The eventual credibility of personal AI agents may depend less on how impressive their demonstrations appear than on whether users can trust them when the situation becomes complicated.
 
(Source:www.cnbctv18.com)

Christopher J. Mitchell
In the same section