Text to Speech

Type some text, pick a voice and language, and get natural-sounding speech you can preview and download.

This one runs on our servers. The speech model is too heavy for a browser, so the text you enter is sent to our GPU servers to be synthesized. Only the text is sent (no file, no account) and the generated audio is removed within two hours. We keep no copy. See the privacy policy.

Text to speech is offline right now

The voice server is not available at the moment. Please try again in a little while.

The voice reads exactly what you type. 0 / 2000
2. Voice and language
Every voice can speak every language. The label next to a voice is just the language it sounds most native in.
3. Download format
The speech is generated once and converted to 44.1 kHz in your browser, so you can switch the format after generating without running the job again.
Checking server status…
Generating on our servers…

Your speech

The audio is in this browser now. Play it above and download it as MP3, FLAC or WAV. We keep no copy on our servers. It lives in this tab only, so download it before you close or reload the page.

To re-encode the file or change its sample rate, use the Format Converter.

What this does

Text to speech turns written text into spoken audio. You type a sentence or a paragraph, choose one of several AI voices and a language, and the model reads it aloud in a natural voice. The result is an audio file you can preview in the browser and download for a video voice-over, a podcast intro, an accessibility narration, a language-learning clip or anywhere else you need a voice without recording one.

Unlike most of the tools here, this one runs on our servers, not in your browser. The speech model is far too large to run on a phone or laptop. So the text is sent to our GPU servers and the audio comes back a moment later. Only the text is sent; nothing is stored.

Voices and languages

There are nine preset voices across several languages, and each one can speak any of the ten supported languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. The language label shown next to a voice is only the accent it sounds most native in; pick the voice you like and the language you want it to read.

MP3, FLAC or WAV

MP3 is smaller and plays everywhere, right for a voice-over or sharing a clip. FLAC is lossless and smaller than WAV, right when the speech is going into a mix. WAV is lossless and uncompressed, the plainest thing to drop onto a DAW track. The model produces the speech once; the file is converted to your chosen format and to 44.1 kHz in your browser, so switching format never re-runs the job, and you can re-encode later with the Format Converter.

Notes, honestly