Voice agents reached operations: what separates a working one from an expensive demo
ElevenLabs says its voice agents run more than 15 million conversations a week, and they now handle returns, insurance actions and bookings. Why voice became the first agentic AI segment with a real production use case, what separates a working agent from a polished demo, and what to settle about recordings and personal data before the pilot starts.
ElevenLabs was valued at roughly 22 billion dollars in an employee share sale, about twice its February mark. The company says its voice agents now run more than 15 million conversations a week, roughly three times the February level. Those numbers come from the company, I have no way to verify them from the outside, and the valuation itself is the least interesting part.
What caught my attention is the breakdown of what those conversations do. This is no longer a speech generation showcase: according to the company, the agents handle returns, insurance actions and bookings. Those are ordinary operational processes that somebody inside a company owns and is measured on. Voice stopped being a shop window and moved into operations. In a field where most agentic AI talk is still about potential, this is the first segment with a visible, repeatable production use case.
What follows is my read on why voice got there first, what separates a working agent from an expensive demo, and the question about personal data you want answered before the pilot rather than after it.
Why voice got there first
I would not explain this with model maturity. Synthesis and recognition did get better, but a support call has four properties that most fashionable agent scenarios lack.
-
The process is short. A return or a booking change fits into minutes. The agent has no long horizon for error to accumulate on, and no chain of ten steps where each one leans on the last. On how poorly agents hold up over long tasks, I wrote separately: long-running tasks and production readiness.
-
The scenario is narrow. A call has a finite set of outcomes: the return is filed, the date is moved, the status is checked, the caller is handed to a person. Narrow scope is exactly where a language model behaves predictably, because there is nowhere for it to wander off to.
-
The result is machine-checkable. An hour after the call you can see whether a return was created, a booking was changed, a ticket was closed. No expert review is needed to tell whether the agent worked. Compare that to an agent writing an analytical memo, where you will argue about the criteria for weeks.
-
The economics already have a denominator. Support knows what an agent-minute costs, how many calls arrive at peak, what it takes to hire and train a person. Any pilot gets compared against a real number instead of a hypothesis. That is usually where I start the AI conversation: questions to ask before the investment.
Put those four together and you get the portrait of a process that survives automation first. Voice itself is almost incidental here; the same logic would work in text. The phone channel simply costs the company the most, so the savings show up sooner.
What separates a working agent from an expensive toy
A voice agent demo almost always looks excellent. It is built to: the script is chosen, the caller is cooperative, the edges are absent. The gap between a demo and a production system sits in five things, and none of them is voice quality.
An honest handoff to a person. The agent has to recognise that it is stuck and pass the call to an operator together with the context: what has been established, what the customer asked for, what the agent already did. A bad handoff means the customer retells the whole story to a human being. That annoys people more than having no agent at all. Build an explicit exit on demand too: when the caller asks for a person, the agent stops negotiating.
Confirmation on irreversible actions. Checking an order status is something the agent can own. Taking money, cancelling a policy, voiding a booking are actions you cannot undo in one move, and they need either an explicit confirmation step or a person in the loop. The line runs along the cost of a mistake rather than the difficulty of the task. I have a separate note on that principle: human in the loop.
Recording and retention that follow the rules. A voice agent produces, by definition, a recording and a transcript of every conversation. That is a new body of sensitive data, and it needs a retention period, access rules, an access log and a deletion procedure. Deal with it after launch and you end up with an archive of customer calls that nobody governs.
Caller identification. Before discussing an order or a policy, the agent needs to know who is on the line. A confident voice is persuasive, and that is precisely the risk: people trust a steady tone. Identification should rest on verifiable factors, such as a code from the app, a call from a known number, or details only the customer holds, and how much is enough gets decided before launch rather than during the first disputed case.
A measured share of calls resolved end to end. The primary metric for a voice agent is how many calls it closed completely, with no human involvement and no callback from the same customer within a day. Everything else clusters around it: handoff rate, average handling time, abandonment, cost per contact. With no such metric you are watching the system rather than running it.
A customer call is personal data, and a voice can be biometrics
I am keeping this as its own section, because it is the part that gets skipped.
Every customer conversation carries personal data: a name, an order number, a delivery address, policy details, sometimes health or financial information. Recording that call and transcribing it is processing of personal data, with everything that follows: a lawful basis, notice to the person that the call is recorded, a retention period, restricted access, deletion on request. A text chatbot creates the same obligations; a phone call just produces more data in a heavier format.
Then the voice-specific part begins. When a voice recording is used to establish who the caller is, it becomes biometric data under a stricter regime almost everywhere. Under GDPR that puts it in the special category of Article 9, which needs its own lawful basis and usually explicit consent. In Russia, 152-FZ treats biometric personal data separately and, as a general rule, requires written consent, with the state biometric system running under its own rules. The practical takeaway is the same in both cases: voice authentication and call recording for service quality are two different projects with different legal weight, and mixing them into one pilot is a mistake.
The second question to put to a vendor is where the recording is physically processed and stored. A voice agent normally means streaming audio into the vendor's cloud, often in another jurisdiction, with recordings and transcripts kept on their side. If you serve Russian customers, the requirement to keep their personal data in databases located in Russia still applies, and cross-border transfer needs its own grounds and regulator notification. Architecturally this comes down to a choice: local deployment, a domestic cloud, or a deliberate narrowing of what leaves your perimeter at all.
The third question is what the vendor does with your conversations. Read the clause about using customer data to improve models carefully: recordings of your customers are your asset and your liability. Before the individual, the controller answers, which means you rather than the model provider. I wrote about building the architecture around that: personal data architecture under GDPR, and about starting from the data flow map: map the flows before adding controls.
How to pick the first process for a pilot
These are the criteria I would screen a candidate against. The process should clear all six, not five of six.
-
High frequency, low variance. Hundreds of similar calls a week. A rare complex case gives you neither statistics nor savings, and hands you every edge condition at once.
-
A short horizon. The task is resolved inside one conversation. Scenarios with waiting, a callback the next day and several participants can wait their turn.
-
A machine-checkable result. After the call an object appears in a system that you can count: a request, a changed record, a closed ticket. That is your success criterion and your metric source.
-
A known cost of error, and a way back. A wrong order status is fixed with an apology. A wrong charge is fixed with a refund, an investigation and reputational damage. Start with the first kind.
-
Clean data on the input side. The agent needs current order status, tariff terms, return rules. When that information is spread across three systems and they disagree, a voice agent will simply narrate your mess faster and more politely.
-
A settled legal footing. Before launch you know the basis for recording, how the person is notified, where audio and transcripts live, for how long, and who can reach them. This gets decided at the start, because retrofitting it means reworking the integration along with it.
One more practical rule: for the first month keep the agent in a mode where a person reviews every recorded call and can step in. That month will hand you the list of real edge cases that nobody can invent in advance.
Honest caveats
-
The ElevenLabs numbers are the company's own claim. The 15 million weekly conversations and the 22 billion dollar valuation come from the company and from an employee share sale. I have no independent verification, and I would not build market conclusions on them. The useful part is the fact itself: the class of processes where voice agents already work has been named out loud.
-
Volume is not quality. A conversation count says nothing about the share of calls the agent resolved, or how often the customer called a human afterwards. Nobody discloses that publicly, and in your own deployment you will have to measure it yourself.
-
Voice costs more than text and fails more visibly. A second and a half of latency goes unnoticed in chat and kills a conversation on the phone. The phone channel adds telephony, streaming recognition, synthesis and barge-in handling, each layer with its own latency and price. If the same process closes in chat, start in chat. On the practical side of voice interfaces I wrote earlier: why voice interfaces matter beyond the phone.
-
I do not sell voice agents. My role is to help choose the process, run the economics, set the legal footing and write the requirements for a vendor. The integration and the telephony are specialist work, and I would rather bring those people in than pretend to be the contractor for everything.
The short version
- ElevenLabs reports 15 million conversations a week and a valuation near 22 billion dollars; what matters more is that the agents are handling returns, insurance actions and bookings.
- Voice went first because of the shape of the process: a short call, a narrow scenario, a machine-checkable result and economics with a ready denominator.
- A working agent differs from a demo in five places: honest handoff, confirmation on irreversible actions, recording and retention rules, caller identification, and a measured share of calls resolved end to end.
- A recorded call is personal data, and a voice used to identify someone is biometric data under a stricter regime. Ask the vendor where the recording is physically processed before the pilot, not after.
- Pick the first process on frequency, short horizon, checkable result, cost of error, clean input data and a settled legal footing.
If you are considering a voice agent in support, or already have a proposal from an integrator on your desk, the conversation worth having is about process choice and the legal footing before the integration rather than after it. Discuss your case through the form on the main page; a first conversation commits you to nothing.