AI Text to Speech

Turn text into natural-sounding speech with Eleven v4, Eleven v4 Turbo or Grok TTS. Dozens of voices, 80+ languages, and inline tags for laughs, pauses and whispers.

Try an example

Narration

Voice: Daniel · Model: Eleven v4

“[calm] Welcome back. Today we're walking through three habits that quietly changed how I work, and the first one takes less than a minute.”

ElevenLabs’ real-time speech model. Half the price of Eleven v4, with the same voices, audio tags and settings.

Price per 1,000 characters: Eleven v4 Turbo 8 credits, Eleven v4 16, Grok TTS 3.

0 / 5,000
Audio tags:

These are examples, not a fixed list: other short directions in square brackets, like [nervous] or [shouting], often work too. To control pronunciation, write the word in IPA between slashes, for example /ˈtoʊmɑːtoʊ/.

21 ElevenLabs preset voices. Press play to hear a short sample before you generate.

Auto-detect works for most text. Pick a language when the text is short or mixes languages, so pronunciation and numbers follow that language.

Voice settings

0.50

Lower is more expressive and varied; higher is steadier and more consistent from take to take.

0.75

How closely the result sticks to the original voice. Higher sounds more like it but can be less natural.

Output

MP3 is small and plays everywhere, including Lip Sync. Opus (.ogg) is smaller still and plays in Chrome, Firefox and Edge; some Safari versions can’t preview it, but the download works everywhere. The price is the same for every format.

Spells out numbers, dates and abbreviations the way they’re read aloud before speaking. Auto lets the model decide; turn it off if your text is already written exactly as it should be read.

Reuse a seed with the same text and settings to get a similar take. An identical result isn’t guaranteed.

Also returns when each character is spoken, as a JSON file you can download next to the audio. Useful for subtitles or syncing animation. Doesn’t change the price.

Cost: 1 credits

Try an example

Narration

Voice: Daniel · Model: Eleven v4

“[calm] Welcome back. Today we're walking through three habits that quietly changed how I work, and the first one takes less than a minute.”

Ad read

Voice: Laura · Model: Eleven v4

“[excited] Tired of recording the same voiceover ten times? [laughs] Same. Type it once, pick a voice, and you're done.”

Storytelling

Voice: Alice · Model: Eleven v4

“The house was quiet. [whispers] Too quiet... And then, somewhere upstairs, a door clicked shut.”

What the text-to-speech models can do

Three text-to-speech models from ElevenLabs and xAI, with expressive delivery written straight into the text.

49 voices across three models

Eleven v4 and Eleven v4 Turbo share 21 ElevenLabs voices, and Grok TTS has 28. Every voice has a short sample you can play before you spend a credit.

80+ languages, or auto-detect

Eleven v4 and Eleven v4 Turbo read 82 languages, from English, Chinese and Spanish to Welsh and Swahili; Grok TTS covers 20. Leave it on auto-detect and the model reads the language from your text.

Laughs, pauses and whispers in the script

With Eleven v4, put audio tags like [whispers], [laughs] or [excited] before the words they should change. Grok TTS uses [laugh], [pause] and [sigh], plus whisper and slow tags; its sample voices on this page use them.

Up to 15,000 characters

Enough for a full script in one go: up to 15,000 characters with Grok TTS, 5,000 with Eleven v4. MP3 by default, with WAV, Opus and raw formats for editing and phone systems.

How to turn text into speech

Three steps, no recording needed.

1

Write or paste your script

Up to 5,000 characters with Eleven v4, or 15,000 with Grok TTS. Add expression tags where you want a laugh, a pause or a whisper.

2

Pick a model, voice and language

Eleven v4 Turbo is selected by default. Press play next to the voice picker to hear any voice first. Leave the language on auto-detect unless your text mixes languages.

3

Generate and download

Your audio appears next to the editor and in My Creations. Download it as an MP3 or in the format you picked.

Frequently asked questions

What the text-to-speech tool does, what it costs, and how it fits with the rest of Veevid.

Give your script a voice

Pick a model and a voice, add a laugh or a pause where it fits, and download the audio.

Start generating