01 Voice & AI
Text to SpeechFree online text to speech with AI voices
Type or paste some text, pick one of 30 English and Spanish voices and download it as natural-sounding speech.
0/600 characters · 0 words · ≈ 0 s of audio at 1.0×
Choose a voice
Reads in EnglishAmerican · female
American · male
British · female
British · male
Spanish · female
Spanish · male
Spanish voices read in Spanish and the rest in English: write the text in the voice's language. Press ▶ to hear a sample without switching voices.
Free · local processing · no account
Spanish voices: Dora, Alex and Santa. Pronunciation and quality vary by voice and language. Up to 600 characters per generation. On computers with a compatible graphics card, an accelerated version (~310 MB, once) downloads after the first audio and later ones take seconds.
First model download: Kokoro · ~90 MB. Processing downloads public model files; your audio is never sent. The first run needs a connection and device storage.
Engines and licensesHow long does it take?
It depends on your device: with a WebGPU-compatible graphics chip it's much faster. Approximate times for a 600-character text (about 40 s of speech):
- Computer with a compatible graphics chip (WebGPU: Chrome, Edge)
- about 30 s, from the second audio on (the first one runs on the processor while the accelerated version downloads, ~310 MB, once)
- Computer without WebGPU (Firefox or unsupported graphics)
- around 1.5 min
- Safari or another browser using a single core
- about 2–2.5 min
- Phone or tablet
- longer than on a computer; it depends a lot on the phone
- The first time, the voice model downloads (about 90 MB). Short texts take proportionally less.
- Everything runs on your device: if you switch tabs or your device goes to sleep, the browser may pause or close the tab and you would have to start over.
About Text to Speech
Text to Speech turns written text into a voiceover with the pauses, intonation and rhythm of a real person. It uses Kokoro, an artificial intelligence model that runs in your own browser: your text is never sent to a server. There are 27 English voices with American or British accents and 3 Spanish voices, and you control the speaking rate. Every result downloads as both WAV and MP3.
What you can do with it
- 01
Voiceovers for your videos
Write the script, choose a voice and get the narration without a microphone, a quiet room or endless retakes.
- 02
Proofread by ear
Hearing what you wrote exposes run-on sentences, repeated words and typos that slip past when you read.
- 03
Study while listening
Turn notes or parts of an article into audio and review them on a walk, at whatever speed feels comfortable.
- 04
English pronunciation
Hear how a word or phrase sounds in an American or British accent and replay it as often as you need.
How it works, in 3 steps
Type or upload your text
Paste the text into the box or upload a .txt file. You'll see the character and word count and roughly how long the audio will be.
Pick a voice and speed
Play each voice's sample before choosing and set the speed between 0.5× and 2×. At 1× it reads at a normal conversational pace.
Generate and download
Click «Generate speech», listen to the result and download it as WAV or MP3. Not quite right? Try another voice or speed and generate again.
The 30 voices
American English, female (11)
Heart, Bella, Sarah, Nicole, Nova, Sky, Alloy, Aoede, Jessica, Kore and River: from warm and friendly to crisp and professional.
American English, male (8)
Adam, Michael, Eric, Liam, Onyx, Echo, Fenrir and Puck, from deep and serious to young and casual, plus a very theatrical one.
British English (8)
Emma, Alice, Isabella and Lily, plus George, Daniel, Lewis and Fable: polished voices that suit education and storytelling.
Spanish (3)
Dora, clear and friendly; Alex, natural and calm; and Santa, deep and festive. Write in Spanish for these voices to sound natural.
Tips for better results
- Spanish voices: Dora, Alex and Santa. Pronunciation and quality vary by voice and language. Up to 600 characters per generation. On computers with a compatible graphics card, an accelerated version (~310 MB, once) downloads after the first audio and later ones take seconds.
- Mind your punctuation: commas create pauses, periods end the sentence and question marks change the intonation.
- Write numbers, acronyms and abbreviations the way you want them read to avoid odd pronunciations.
- If a name or technical term sounds wrong, spell it the way it's pronounced or rephrase the sentence.
- Try the same text with two or three voices: the best fit is often not the one you'd have picked first.
- For texts over 600 characters, generate several parts and join them with Audio Joiner.
Frequently asked questions
Is this text to speech really free?
Yes, with no sign-up and no usage limit. The voice model downloads once and runs in your browser, so there are no servers to pay for each audio.
Which accents are available?
American and British English, with 27 voices between them, plus 3 Spanish voices. Write the text in the language of the voice you choose.
How much text can I convert at once?
Up to 600 characters per generation, about 40 seconds of speech. For longer texts, split them into parts and join them with Audio Joiner.
Should I download WAV or MP3?
You get both. WAV keeps full quality and is best if you'll edit the voice; MP3, at 192 kbps, is much smaller and perfect for publishing or sharing.
Does changing the speed make the voice worse?
No. Speed is applied while the voice is generated, not by speeding up playback, so it doesn't distort or change pitch at 0.5× or at 2×.
Is my text sent to a server?
No. The speech is generated on your device with the Kokoro model; the only thing downloaded from the internet is the model itself, the first time.
Can it speak with my own voice?
That's what Voice Cloning does, imitating a voice from a sample a few seconds long. It's experimental and currently reads English only.