Skip to content
Rush Commerce
AI & Automation4 min read

Eleven v4 Turbo hits 150ms — voice agents got usable

ElevenLabs shipped Eleven v4 and v4 Turbo with ~150ms time to first speech, 90+ languages, and 10-second voice cloning. What it changes for phone automation.

The thing that made AI phone agents feel wrong was never the voice. It was the pause. On September 28 ElevenLabs shipped Eleven v4 and Eleven v4 Turbo, and v4 Turbo reports a median ~150ms time to first speech with ~100ms median inference latency. Human conversational turn-taking runs in roughly that range. For anyone who shelved a voice agent project because callers kept talking over dead air, the technical objection just got weaker.

What actually happened

Two models, one split. Eleven v4 is the expressive one, aimed at pre-recorded audio. Eleven v4 Turbo trades some of that for speed and is explicitly built for live agents. ElevenLabs publishes v4 Turbo at ~150ms median time to first speech and puts competing models between 262ms and 814ms on its own benchmark — treat a vendor's own comparison as a vendor's own comparison, but the absolute number is the part you can test yourself.

Language coverage moves from roughly 70 to over 90 languages, with the company calling out the largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. A single voice can speak any supported language while holding its speaker identity.

Instant voice cloning now needs 10 seconds of audio. Professional Voice Clones are supported for high-fidelity work. Inline audio tags carry over from v3 and got stricter adherence — [laughs], [sleepy drowsy voice], [said angrily in French accent], even environmental cues like [light rain] — and you can stack tags and have the model follow the sequence. IPA phoneme support improved, which matters more than it sounds if your catalog is full of product names no model has seen.

The models are live across ElevenAgents, ElevenCreative and the API. ElevenLabs did not publish v4 pricing in the announcement, and says 55% of its revenue now comes from enterprise customers, with an annualized run rate that moved from about $330M to over $600M inside the year.

Why faster voice AI matters for your business

Latency was the whole product. A 700ms gap before a reply reads as a bad connection or a bored employee. Around 150ms reads as a person. That single number decides whether callers stay on the line, and it is the first thing to measure in a pilot — end to end, on your network, with your telephony provider in the path, not from a vendor slide.

Turbo means you stop pretending the LLM is instant. v4 Turbo can start generating audio as soon as the language model behind it starts producing tokens. That is an architecture change, not a setting: your agent needs to stream, and any synchronous step you inserted — a CRM lookup, a stock check, a policy retrieval — now sits in the caller's ear as silence. Move those calls behind the first sentence or cache them.

10-second cloning is a consent problem before it is a feature. You can now clone a staff member's voice from a voicemail greeting. Get written permission, keep the recording and the authorization together, and decide up front whether your brand voice is a real employee or a synthetic one. One of those follows you when they quit.

90 languages is a market decision disguised as a model release. If a fifth of your inbound is Spanish or Cantonese and you handle it with a callback queue, the cost of covering it in-line just dropped. That is revenue, not tooling.

Disclose the bot. The model being good enough to pass is exactly why you say it is a bot at the top of the call. Regulation is tightening in that direction and customers punish the discovery harder than the disclosure.

Key takeaways

  • ElevenLabs released Eleven v4 and v4 Turbo on September 28, 2026 across ElevenAgents, ElevenCreative and the API
  • v4 Turbo reports ~150ms median time to first speech and ~100ms median inference latency; v4 favors expressiveness for pre-recorded audio
  • Language support grew from ~70 to over 90, with the biggest gains cited in Japanese, Brazilian Portuguese, Mandarin and Cantonese
  • Instant voice cloning needs 10 seconds of audio — treat voice consent and authorization records as a requirement, not paperwork
  • Pricing was not published in the announcement; benchmark latency yourself, end to end, with your telephony provider in the path

A voice agent is a phone system problem wearing an AI costume. We build call automation where the transcript, the CRM write and the escalation path live in systems you own, so swapping the speech vendor is a config change instead of a rebuild. See how we wire voice into your stack, or tell us what your callers are waiting on hold for.

Sources: Eleven v4, ElevenLabs, ElevenLabs' new v4 speech model supports more expression control and 90 languages, TechCrunch.

  • #voice-ai
  • #elevenlabs
  • #ai-agents
  • #customer-support
  • #latency
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.