03 Voice & AI
Voice CloningClone a voice with AI in your browser
Record a few seconds of your voice and have it read any English text you type, with the same timbre.
“I like walking by the river as the evening falls. Is there anything better than a mild, quiet day? Most people would say no, though some prefer the cool nights of autumn.”
Speak clearly at your own pace, in a quiet room and close to the microphone. Use an external or headset mic if you have one: the result copies what it hears.
Recording time and limit:0:00 / 0:20
0/200 characters · 0 words · ≈ 0 s of audio at 1.0×
Free · local processing · no account
Experimental, English only: the multilingual variant is not yet compatible with this browser engine. Requires WebGPU and substantial memory. Use a clean 5–15 s sample (longer ones are trimmed to the first 15 s) and up to 200 characters. Only use a voice you have permission to use.
First model download: Chatterbox · >1 GB. Processing downloads public model files; your audio is never sent. The first run needs a connection and device storage.
Engines and licensesHow long does it take?
It depends on your device: with a WebGPU-compatible graphics chip it's much faster. Approximate times for a short sentence:
- Computer with a compatible graphics chip (WebGPU: Chrome, Edge)
- several minutes; the first time it downloads the model, around 1.5 GB
- Computer without WebGPU (Firefox or unsupported graphics)
- not available: it needs WebGPU
- Safari or another browser using a single core
- not available: it needs WebGPU
- Phone or tablet
- not available on most phones
- Everything runs on your device: if you switch tabs or your device goes to sleep, the browser may pause or close the tab and you would have to start over.
About Voice Cloning
Voice Cloning imitates the timbre of a voice from a 5 to 15-second sample and uses it to read the text you write. It runs Chatterbox, an AI model that works in your browser thanks to WebGPU, so the sample is never uploaded to a server. It's an experimental feature: for now it only reads English and needs a computer with compatible graphics. Only clone your own voice or the voice of someone who has given you permission.
What you can do with it
- 01
English voiceovers in your voice
Record the sample once and generate lines for your videos or presentations by typing the script, without stepping back up to the mic.
- 02
Fix a single line
If you stumbled over a sentence in an English recording, generate the correct version in your voice and swap it in without redoing the whole take.
- 03
Hear how you'd sound
Listen to your voice reading texts you never recorded and try voice cloning without sending your voice to any company.
- 04
Audio prototypes
Create placeholder voices for an animation, a game or a demo before recording the final version.
How it works, in 3 steps
Provide the sample
Record about 15 seconds reading the sample text or upload clean audio of the voice. If it's longer, the first 15 seconds are used.
Type the English text
Write what you want it to say, up to 200 characters. Adjust the speed and the expressiveness, from restrained to intense.
Confirm permission and generate
Tick the box confirming it's your voice or you have permission, generate, and download the result as WAV or MP3.
What you need
A browser with WebGPU
A recent Chrome or Edge on a computer with compatible graphics. Without WebGPU, the tool can't run.
Memory and patience
The first time it downloads a model of more than 1 GB, and generation can take several minutes. After that the model stays saved in your browser.
A clean sample
5 to 15 seconds of a single voice, with no music or background noise. The clearer it is, the closer the result.
English text
The model that runs in the browser only speaks English. For Spanish voices, use Text to Speech.
Tips for better results
- Experimental, English only: the multilingual variant is not yet compatible with this browser engine. Requires WebGPU and substantial memory. Use a clean 5–15 s sample (longer ones are trimmed to the first 15 s) and up to 200 characters. Only use a voice you have permission to use.
- Record in a quiet room and, if you can, with an external mic or a headset mic.
- Speak clearly at your normal pace: the clone imitates what it hears in the sample, hesitations included.
- If the sample has background noise, clean it first with Noise Remover.
- Raise the expressiveness for a more emotional tone and lower it for a neutral read.
- Only clone your own voice or someone who has given you explicit permission: doing it without consent may be illegal.
Frequently asked questions
How much audio do I need?
5 to 15 seconds of clear speech. If you record inside the tool, there's a sample text ready for you to read aloud.
Can it speak other languages?
Not yet. The multilingual version of the model doesn't run in the browser yet, so the text has to be in English. For Spanish, use the Text to Speech voices.
Why does it need WebGPU?
The model is large and only runs at a reasonable speed on the computer's graphics chip. Chrome and Edge use it through WebGPU.
Can I clone someone else's voice?
Only your own voice or with the owner's explicit permission. Cloning someone's voice without consent can violate their rights and the law.
Is my voice uploaded to a server?
No. The sample and the text are processed on your device; only the model is downloaded from the internet.
How is it different from Text to Speech?
Text to Speech uses 30 ready-made voices in English and Spanish. Voice Cloning builds the voice from the sample you provide, so the result sounds like that person.