AI Speech to Text

Convert audio to text in minutes. Upload a recording and get an accurate transcript with speaker labels and timestamps, ready to copy or download as TXT, SRT or VTT.

ElevenLabs Scribe V2

MP3, WAV, M4A, OGG, FLAC, AAC · up to 100MB and 60 minutes per file

Detect who is speaking and split the transcript by speaker.

Mark laughter, applause and music, e.g. [applause].

0 / 100

Separate with commas. Helps the model spell names, brands and jargon correctly. Costs 3 credits per minute instead of 2.

2 credits per minute of audio

Example transcript

0:00Speaker 1

Hey Maya, did you get a chance to look at the launch plan for Veevid's new speech-to-text page?

0:05Speaker 2

I did. Honestly, the keyword research surprised me. Audio-to-text has less volume, but the intent is much closer to what we're building.

0:13Speaker 1

Right. People searching voice-to-text mostly want dictation on their iPhone or Android, not file uploads.

Copy text TXT · SRT · VTT 90+ languages

Real output from a two-person test recording. Your transcript will show here with timestamps and speaker labels.

Audio to text, done properly

Upload a recording and get a clean, readable transcript: who said what, when they said it, and every name spelled the way you want.

Accurate in 90+ languages

Powered by ElevenLabs Scribe V2, one of the most accurate speech recognition models on independent word-error-rate benchmarks. It detects the language automatically, so you do not need to pick one.

Speaker labels

Interviews, podcasts and meetings come back split by speaker, with each turn marked Speaker 1, Speaker 2 and so on, plus the time it starts.

Timestamps and subtitle files

Every word is timed, so you can download ready-made SRT or VTT subtitles for your video, or a plain TXT transcript for notes and articles.

Names and jargon spelled right

Add the names, brands and technical terms in your recording and the model uses them. In our tests, a brand name it got wrong on its own came out right once we listed it.

Sound tags

Laughter, applause and music are marked in the transcript, like [applause], so captions tell the whole story. Turn it off if you only want the words.

Fast, even for long files

An 18-minute speech took about 15 seconds to transcribe in our tests. Files up to 60 minutes are supported, and every transcript is saved to My Creations.

How to convert audio to text

Three steps from recording to transcript. No software to install.

1

Upload your audio

Drop in an MP3, WAV, M4A, OGG, FLAC or AAC file up to 60 minutes long, or paste a link to one. The price shows as soon as the length is read.

2

Choose your options

Keep speaker labels and sound tags on or off, and list any names or terms that must be spelled exactly.

3

Transcribe and download

Click Transcribe. Copy the text or download it as TXT, SRT or VTT. The transcript is also saved in My Creations.

Speech to text FAQ

How the audio to text converter works, what it costs and what it can handle.

Turn your next recording into text

Upload an audio file and get a transcript with speakers, timestamps and subtitles in about a minute.

Transcribe audio now