Eleven v4: Our most expressive text-to-speech AI model yet
The text-to-speech models support over 90 languages, inline emotion prompting, and a median time to first speech of ~150ms on the Turbo variant.
- Supports natural language inline audio tags for emotions, accents, and sound effects, alongside improved IPA phoneme support.
- Eleven v4 Turbo achieves a median inference latency of ~100ms for conversational agent workflows.
- Instant Voice Clones can replicate voices using 10 seconds of audio while maintaining speaker identity across language switches.
- Both models are available immediately through ElevenAgents, ElevenCreative, and the ElevenAPI.
Developers and creators get more expressive multi-speaker voice synthesis with lower latency and fine-grained natural language delivery controls.

Sources
Read this as text
Back to the AI news