ElevenLabs Unveils v4 Voice Model With 90 Languages & Advanced Expression Control

ElevenLabs has rolled out its next-generation speech models—v4 and v4 Turbo—bringing major upgrades in language support, expressive voice control, and latency. These releases mark a significant leap over its v3 systems, aiming to streamline voice cloning and make generated speech sound more natural, especially in interactive settings.

What’s New

The v4 architecture introduces the ability to clone a voice using only 10 seconds of audio. It also sustains voice identity over longer passages of text and enhances expressive reading by tracking text context. New inline tag functionality allows users to layer multiple expressive cues—for example, combining tone and mood—and ensures the model follows the tagged sequence.

Language coverage has jumped significantly. While the previous generation handled about 70 languages, v4 now supports more than 90. The most noticeable improvements come in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, where quality enhancements are most pronounced.

Enterprise & Performance Gains

ElevenLabs reports that enterprise clients now make up over 55% of its business, underscoring demand for high-performance voice agents. With lower latency and the ability to start audio output as soon as the underlying large language model (LLM) begins producing results, v4 is built for more fluid conversational experiences. The model also better handles challenging dynamics like customer escalation and holds to improve issue resolution.

Financial Trajectory & Company Growth

Earlier this year, ElevenLabs raised $500 million in a funding round led by Sequoia, placing its valuation at $11 billion. Speculation is already swirling around a potential subsequent round that could hike its valuation to $22 billion. The company’s revenue run rate has surged from roughly $330 million at the start of the year to over $600 million.

To keep up with global demand and scale, ElevenLabs has expanded hiring, boosting its workforce to over 800 employees spread across markets including India, Europe, and Brazil. Internally, leadership is discussing an initial public offering sometime in “the next years,” though no exact timeline has been confirmed.

Competition in expressive speech AI is intensifying. Startups like Cartesia, Deepgram, Fish Audio, Boson, and WellSaid Labs are all pushing boundaries, while major players such as Google and OpenAI continue evolving their own voice model offerings.

Why It Matters: Voice is rapidly becoming a frontline medium in AI—powering everything from virtual assistants to automated agents. Systems that can deliver natural, expressive speech in more languages help level the playing field globally and make voice interfaces more inclusive and effective.

What to Watch: The rollout of v4 raises questions about ethical voice cloning, transparency around data and model behavior, and how license and usage policies will evolve. Also, whether ElevenLabs can convert its momentum into a successful public market debut will shape its competitive landscape.