9 min readTranscription, Subtitles, Whisper, Tutorial
How to transcribe audio to text for free and get SRT subtitles
Transcribe interviews, lectures or meetings with Whisper in your browser: which accuracy to pick, how to improve results and how to use SRT subtitles.
Transcribing by hand is slow: an hour-long interview can take four or five hours of typing, pausing and rewinding. Automatic speech recognition turns that work into reviewing and correcting. The key is knowing what to expect from it, preparing the audio well and choosing the right setting for each recording.
What Whisper is
Whisper is the speech recognition model OpenAI released openly in 2022. It was trained on about 680,000 hours of audio in dozens of languages, which is why it handles accents, moderate noise and varied vocabulary far better than earlier systems. Our tool runs it inside your browser: the audio is never uploaded to a server, which matters when you transcribe confidential interviews, work meetings or lectures with student data.
The model works in 30-second chunks and shows its progress, so you can transcribe long recordings. It recognizes 33 languages in the selector and can also detect the language on its own.
Which accuracy to choose
| Accuracy | Download | When to use it |
|---|---|---|
| Fast | ≈ 40 MB | Clear speech, no noise; phones and modest computers |
| Balanced | ≈ 80 MB | Almost everything: good accuracy without a long wait |
| High | ≈ 250 MB | Strong accents, technical vocabulary, hard audio |
In our tests on a current laptop, three minutes of audio at Balanced accuracy are transcribed in under a minute. On phones and older computers it takes considerably longer, and the model only downloads the first time.
Step by step
- Get the audio. Upload an MP3, WAV, M4A, OGG or FLAC, or click Record to speak into the microphone (up to 10 minutes). If you have a video, extract its sound first with Video to Audio.
- Choose the language. If you know it, select it: it's more reliable than auto-detection, especially on short recordings.
- Choose the accuracy using the table above.
- Transcribe and review. Read the text with the audio at hand and fix names, acronyms and technical terms.
- Export. Copy the text, download it as TXT or download the SRT to subtitle a video.
How to improve the transcript
- Clean the noise first. A fan, traffic or electrical hum increase errors. Run the recording through the noise remover in Voice mode before transcribing.
- Trim long silences. With a lot of silence, the model sometimes 'invents' sentences nobody said. Remove them with the trimmer.
- Avoid background music. If there is some, the voice should sit clearly above it.
- One language per recording. If two are mixed, choose the one spoken most.
- Record close to the microphone. A hand's width away is understood far better than two metres away, even if the recording sounds quieter.
What an SRT file is and how to use it
SRT is the most widespread subtitle format. It's a text file where each block has a number, a start and end time and the line. For example, block 1 might run 00:00:01,000 --> 00:00:04,200 with the text 'Hello and welcome to the episode'. Almost every video editor and platform accepts it: you import it and the subtitles appear in sync.
- Review before importing. Fixing a word in the SRT is easy in any text editor; don't touch the timing lines.
- Short lines. If a sentence is very long, split it into two blocks so it can be read in time.
- Save as UTF-8. That way accents and special characters display correctly in any player.
What to expect from the accuracy
With one person speaking clearly and little noise, the transcript is very faithful and will only need touch-ups. Errors increase with strong accents, people talking over each other, loud music or poor recording quality. What most often fails is proper names, acronyms and industry jargon: always check those parts, especially if you're going to quote someone.
FAQ
Does it tell speakers apart?
No. The transcript is continuous text with timestamps; if you need to show who's speaking, add it while reviewing.
Can it get the lyrics of a song?
With background music, accuracy drops a lot. It works much better if you first separate the voice with the Vocal Remover and transcribe only the acapella, although singing is still harder than speech.
Can I transcribe a WhatsApp or iPhone voice note?
Yes. Voice notes are usually M4A or OGG, and both formats are accepted directly. If some program won't open them, convert them first with the audio converter.
Keep reading
9 min read
Free AI text to speech: how to get a voiceover that sounds natural
How modern text to speech works, how to write a script that sounds human, which voice and speed to choose and how to build long voiceovers.
9 min read
How to remove background noise from audio without ruining the voice
Which noises can be removed and which can't, how to choose the mode and strength, and the mistakes that leave a voice sounding metallic or canned.
8 min read
How to make a ringtone from any song (iPhone and Android)
Cut the perfect clip, add a fade and set it as your ringtone on Android or iPhone. Length, format and the steps for each system.