everyaudiotool

Free AI text to speech: how to get a voiceover that sounds natural

How modern text to speech works, how to write a script that sounds human, which voice and speed to choose and how to build long voiceovers.

A few years ago you could spot a synthetic voice instantly: flat, pausing in odd places and with an intonation that rose and fell for no reason. Today's text-to-speech models are neural networks trained on thousands of hours of real speech, and they learn something the old phonetic rules never captured: prosody, the rhythm, pauses and melody a person uses to say a sentence. That's why a good synthetic voiceover can now pass for a microphone recording.

Which model we use and why it runs in your browser

Our text to speech uses Kokoro, an open model of about 82 million parameters released under the Apache 2.0 license. It's small next to the big cloud services, but very well trained, which makes it possible to run inside the browser: the first time it downloads (about 90 MB in its processor version) and then it stays in the cache. Your text is never sent to a server, which matters if you're converting unpublished scripts, personal notes or work documents.

There are 30 voices: 27 in English, with American or British accents, and 3 in Spanish. If your computer has WebGPU-compatible graphics, the tool downloads a heavier version of the model in the background that generates noticeably faster for the following clips.

Step by step

  1. Prepare the text. Type or paste it into the box; you can also upload a .txt file. You'll see the characters, words and an estimate of the length.
  2. Listen to the samples. Each voice has an example. Write in the language of the voice you pick: an English voice reading Spanish sounds forced.
  3. Set the speed. Between 0.5× and 2×. Speed is applied while generating the voice, not by speeding up playback, so it doesn't distort.
  4. Generate and review. Listen to the full result before downloading. If a word sounds odd, fix it in the text and generate again.
  5. Download. MP3 at 192 kbps to publish or share; WAV if you'll keep editing the voice.

How to write so it sounds human

The model reads what it sees. The difference between a passable voiceover and one that sounds natural is almost always in the script, not the voice. These are the rules that make the biggest difference:

  • Punctuation is your acting direction. A comma creates a short pause, a period closes the idea and lowers the intonation, and a question mark raises it at the end. If a sentence sounds rushed, add a comma; if it sounds choppy, remove one.
  • Short sentences. A 40-word sentence is hard to read even for a person. Split long ideas into two or three sentences.
  • Write numbers the way you want to hear them. '1/2' can be read several ways; 'half' only one.
  • Expand abbreviations and unclear acronyms. 'Dr.' is safer as 'doctor'; if you want an acronym spelled out, separate the letters with hyphens.
  • Avoid web addresses and symbols. A URL read aloud helps nobody: say 'on our website' and put it in the video description.
  • Tricky proper names: spell them the way they're pronounced, or rephrase the sentence if the model stumbles.
You writeBetter asWhy
1/2 cuphalf a cupAvoids 'one slash two'
Dr. Smithdoctor SmithNo ambiguity
10-12 hten to twelve hoursRanges with hyphens confuse it
Amazing!!!Amazing!Repeating adds no emphasis

Which voice and speed to choose

No voice is better than another: some voices fit a piece of content. Clear, calm voices work for tutorials and training; more energetic ones for short social videos. A useful trick is to generate the same paragraph with two or three voices and listen to them back to back: the one you end up choosing is rarely the one you'd have picked from the description.

  • 0.9×: explanations with figures, languages the listener is learning, content for older audiences.
  • 1×: the pace of a normal conversation. The starting point for almost everything.
  • 1.1–1.2×: short videos and ads, where every second counts.
  • 1.5× or more: only for reviewing your own texts; for an audience it's tiring.

Long texts: how to build a full voiceover

Each generation takes up to 600 characters, about 40 seconds of speech. For a script several minutes long, split it by whole paragraphs (never mid-sentence), generate each part with the same voice and speed, and join them with the audio joiner using a 0-second crossfade, which is right for speech. If one part sounds quieter than another, run the result through the volume booster in Normalize mode to bring it all to -16 LUFS.

What people use it for

  • Video narration without a microphone, a quiet room or endless retakes.
  • Proofreading by ear: hearing what you wrote exposes long sentences, repetitions and typos.
  • Accessibility: turning notices, instructions or materials into audio for people who prefer listening.
  • Prototypes: testing an ad or podcast script before recording it with a person.

FAQ

Why is the first time slower?

Because the model downloads. From then on it's stored in your browser and a short sentence is generated in a few seconds; on computers without compatible graphics it takes a little longer.

Can I use the audio in my videos?

Kokoro is released under the Apache 2.0 license, which allows commercial use of the model. Even so, each platform has its own rules on AI-generated content: check them before publishing.

What about the other way, speech to text?

That's what Speech to Text is for. We explain how to get the most out of it in our guide to transcribing audio.

  1. How to transcribe audio to text for free and get SRT subtitles

    Transcribe interviews, lectures or meetings with Whisper in your browser: which accuracy to pick, how to improve results and how to use SRT subtitles.

  2. How to remove background noise from audio without ruining the voice

    Which noises can be removed and which can't, how to choose the mode and strength, and the mistakes that leave a voice sounding metallic or canned.

  3. How to make a ringtone from any song (iPhone and Android)

    Cut the perfect clip, add a fade and set it as your ringtone on Android or iPhone. Length, format and the steps for each system.

← All articles