Eleven v4 Text to Speech

Turn a script into expressive speech with ElevenLabs Eleven v4. Direct the delivery with audio tags like [whispers] and [laughs], pick a voice and language, and download the audio.

Try an example

Narration

Voice: Daniel · Model: Eleven v4

“[calm] Welcome back. Today we're walking through three habits that quietly changed how I work, and the first one takes less than a minute.”

ElevenLabs’ most expressive speech model, for narration, audiobooks and character voices.

Price per 1,000 characters: Eleven v4 Turbo 8 credits, Eleven v4 16, Grok TTS 3.

0 / 5,000
Audio tags:

These are examples, not a fixed list: other short directions in square brackets, like [nervous] or [shouting], often work too. To control pronunciation, write the word in IPA between slashes, for example /ˈtoʊmɑːtoʊ/.

21 ElevenLabs preset voices. Press play to hear a short sample before you generate.

Auto-detect works for most text. Pick a language when the text is short or mixes languages, so pronunciation and numbers follow that language.

Voice settings

0.50

Lower is more expressive and varied; higher is steadier and more consistent from take to take.

0.75

How closely the result sticks to the original voice. Higher sounds more like it but can be less natural.

Output

MP3 is small and plays everywhere, including Lip Sync. Opus (.ogg) is smaller still and plays in Chrome, Firefox and Edge; some Safari versions can’t preview it, but the download works everywhere. The price is the same for every format.

Spells out numbers, dates and abbreviations the way they’re read aloud before speaking. Auto lets the model decide; turn it off if your text is already written exactly as it should be read.

Reuse a seed with the same text and settings to get a similar take. An identical result isn’t guaranteed.

Also returns when each character is spoken, as a JSON file you can download next to the audio. Useful for subtitles or syncing animation. Doesn’t change the price.

Cost: 1 credit

Try an example

Narration

Voice: Daniel · Model: Eleven v4

“[calm] Welcome back. Today we're walking through three habits that quietly changed how I work, and the first one takes less than a minute.”

Ad read

Voice: Laura · Model: Eleven v4

“[excited] Tired of recording the same voiceover ten times? [laughs] Same. Type it once, pick a voice, and you're done.”

Storytelling

Voice: Alice · Model: Eleven v4

“The house was quiet. [whispers] Too quiet... And then, somewhere upstairs, a door clicked shut.”

Hear Eleven v4

Real generations made with the generator above. Every script is shown exactly as it was sent, audio tags included, with the model and voice that read it.

A whisper, then nerves

[whispers] drops the voice to a whisper, and [nervous] changes how the next question lands.

Model: Eleven v4Voice: Charlotte

[whispers] Keep your voice down. They are right outside the door. [nervous] Did you lock it? Please tell me you locked it.

A laugh in the middle of a line

[laughs] is performed between two sentences of the joke, not read out.

Model: Eleven v4Voice: Roger

So I told him the meeting was at nine. [laughs] He showed up at nine, all right. Nine in the evening.

Sad to excited in one take

Three tags switch the mood twice in under ten seconds of audio.

Model: Eleven v4Voice: Jessica

[sad] I really thought we had lost it all this year. [excited] And then the call came in. We got it! [laughs] We actually got it!

A Chinese script

Language set to Chinese. The audio tags stay in English square brackets inside the Chinese text.

Model: Eleven v4Voice: LilyLanguage: Chinese (zh)

[calm] 欢迎收听今天的节目。[whispers] 先偷偷告诉你一个小秘密:明天会有一个惊喜。

Same line: Eleven v4 vs Eleven v4 Turbo

Same script, same voice and default settings, generated once with each model. Listen to both before you choose.

[excited] Type a line, pick a voice, and hear it come to life in seconds. [laughs] No microphone needed.

Model: Eleven v4Voice: Brian
Model: Eleven v4 TurboVoice: Brian

What Eleven v4 can do

ElevenLabs' most expressive speech model, with all of its generation settings available on Veevid.

Direct the delivery with audio tags

Put a tag in square brackets before the words it should change, such as [whispers], [laughs], [sighs], [excited], [sad] or [sarcastic]. Tags are performed rather than read out, and you can switch emotion mid-script.

82 languages

From English, Chinese, Cantonese and Spanish to Hindi, Arabic, Welsh and Swahili. Leave it on auto-detect, or set the language when the text is short or mixes languages.

21 preset voices

Rachel, Aria, Roger, Charlotte, Brian, Lily and the rest of the ElevenLabs preset voices, shared by Eleven v4 and Eleven v4 Turbo. Press play next to the voice picker to hear any of them for free.

Stability, similarity and seed

Lower stability for a more expressive, varied read, or raise it for consistent takes. Similarity sets how closely the output sticks to the voice, and a seed helps you get a similar take again.

19 output formats

MP3 up to 192 kbps, Opus in an .ogg file, raw PCM from 8 to 48 kHz, and 8 kHz μ-law or A-law for phone systems. The price is the same for every format.

Character-level timestamps

Turn on timestamps to get the start and end time of every character as a JSON file next to the audio, ready for subtitles or syncing animation.

Eleven v4 or Eleven v4 Turbo?

Both use the same voices, audio tags, settings and 5,000-character limit. The difference is expressiveness, speed and price.

Eleven v4

16 credits / 1,000 characters

ElevenLabs' most expressive speech model, with the widest emotional range of the two.

Best for: narration, audiobooks, character voices and anything where the performance matters most.

Eleven v4 Turbo

8 credits / 1,000 characters

ElevenLabs' real-time version of v4, built for low latency, at half the price. On Veevid both return a finished audio file, so the practical difference is the delivery and the cost.

Best for: drafts, long scripts, voice-agent lines and high-volume work.

How to use Eleven v4

Three steps, no microphone needed.

1

Write your script

Paste up to 5,000 characters and add audio tags like [whispers] or [laughs] where the delivery should change.

2

Pick a model, voice and settings

Eleven v4 is selected by default; switch to Turbo for the lower price. Choose a voice and language, then adjust stability, similarity, output format or timestamps if you need to.

3

Generate and download

The audio appears next to the editor and in My Creations. Download it in the format you picked, or send an MP3 straight to Lip Sync.

Eleven v4 FAQ

Common questions about using Eleven v4 on Veevid.

Eleven v4 is ElevenLabs' latest text-to-speech model family, which ElevenLabs describes as its most emotive speech model. On Veevid you can use Eleven v4 and Eleven v4 Turbo from this page or from AI Text to Speech, with 21 preset voices and 82 languages.

Eleven v4 is the more expressive model; Turbo is the real-time version built for low latency. They share the same voices, audio tags, settings and 5,000-character limit. Eleven v4 costs 16 credits per 1,000 characters and Turbo 8.

Write short directions in square brackets right before the words they should affect, for example [whispers], [laughs], [sighs], [excited], [sad], [nervous] or [sarcastic]. There is no fixed list, so other short directions often work too. If a tag has little effect, try a more common wording or lower the stability setting, which lets the model deliver more expressively.

Stability goes from 0 to 1 (default 0.5). Lower values allow a more expressive, varied delivery; higher values make takes more consistent. Similarity also goes from 0 to 1 (default 0.75) and controls how closely the output follows the chosen voice; higher values sound more like it but can be less natural.

82 languages on Veevid, including English, Chinese, Cantonese, Japanese, Korean, Hindi, Arabic, Spanish, Portuguese, French and German. Leave the language on auto-detect for most text, or set it when the script is short or mixes languages, so pronunciation and numbers follow that language.

Write the word in IPA between forward slashes, for example /ˈtoʊmɑːtoʊ/, and the model uses that pronunciation. Text normalization (auto, on or off) controls whether numbers, dates and abbreviations are spelled out before they are spoken.

Yes. Turn on character timestamps before generating and you can download a JSON file with the start and end time of every character, grouped by word. It is available next to the result and in My Creations, and it doesn't change the price.

Eleven v4 costs 16 credits and Eleven v4 Turbo 8 credits per 1,000 characters, rounded up, with a minimum of 1 credit per generation. Characters are counted after trimming leading and trailing spaces, and audio tags count toward the length. Credits are refunded if a generation fails.

More audio tools

Keep going with the rest of Veevid's audio and video tools.

Give your script a voice with Eleven v4

Write a line, add a tag, and hear it in seconds.

Start generating