everyaudiotool

02 Voice & AI

Speech to TextTranscribe audio to text online, free

Turn interviews, lectures or meetings into editable text and export SRT subtitles for your videos.

Drop your audio here

MP3, WAV, M4A, OGG, FLAC, WEBM · up to 100 MB and 10 minutes

Choose the spoken language: the transcript will be more accurate.

Accuracy
Free · local processing · no account

Review the transcript: it can make mistakes or invent words, especially with silence or noise.

First model download: Whisper multilingüe / multilingual · 40–250 MB. Processing downloads public model files; your audio is never sent. The first run needs a connection and device storage.

Engines and licenses

How long does it take?

It depends on your device: with a WebGPU-compatible graphics chip it's much faster. Approximate times for 10 minutes of audio at «Balanced» accuracy:

Computer with a compatible graphics chip (WebGPU: Chrome, Edge)
about 4 min (transcription runs on the processor; the graphics chip doesn't speed it up)
Computer without WebGPU (Firefox or unsupported graphics)
about 4 min
Safari or another browser using a single core
about 13 min
Phone or tablet
considerably longer; best with short files
  • «Fast» takes about half as long and «High» about three times longer. The first time, the model downloads (40–250 MB depending on accuracy).
  • Everything runs on your device: if you switch tabs or your device goes to sleep, the browser may pause or close the tab and you would have to start over.

About Speech to Text

Speech to Text listens to a recording, or to what you say into the microphone, and turns it into written text. It runs Whisper, OpenAI's speech recognition model, right in your browser: your audio never leaves your device. It recognizes 33 languages, including English, Spanish, French, German and Chinese, or detects the language on its own. Copy the transcript, download it as TXT or as SRT subtitles with the timing of every line.

What you can do with it

  • 01

    Subtitles for your videos

    Download the SRT with the timings already in place and load it into your editor or the platform where you publish.

  • 02

    Lecture notes

    Record the class, transcribe it and search the text for what was explained instead of scrubbing through the audio.

  • 03

    Accurate interview quotes

    With timestamps you can find each statement in the recording and check the quote word for word.

  • 04

    Voice memos and meetings

    Turn voice notes or a meeting recording into text you can archive, search or share with people who missed it.

How it works, in 3 steps

  1. Record or upload audio

    Click «Record» to speak into the microphone (up to 10 minutes) or upload an MP3, WAV, M4A, OGG, FLAC or other common audio file.

  2. Choose language and accuracy

    Pick the spoken language or leave it on auto-detect, and choose the accuracy: Fast, Balanced or High.

  3. Transcribe and export

    Click «Transcribe», review the text, show or hide timestamps and download it as TXT or SRT, or copy it to the clipboard.

Accuracy and formats

  • Fast (~40 MB)

    The lightest model. Works well with clear speech and little noise, and it's the quickest option on phones and modest computers.

  • Balanced (~80 MB)

    The recommended choice for almost everything: good accuracy without a long wait.

  • High (~250 MB)

    The most accurate, for strong accents, technical vocabulary or difficult audio. It's also the slowest.

  • Plain text (TXT)

    The clean transcript, ready to paste into a document, an email or an article.

  • Subtitles (SRT)

    Each segment with its start and end time, in the subtitle format almost every video editor and platform accepts.

Tips for better results

  • Review the transcript: it can make mistakes or invent words, especially with silence or noise.
  • The cleaner the recording, the more faithful the text: if there's background noise, run it through Noise Remover first.
  • Trim long silences or parts you don't need with Audio Trimmer before transcribing.
  • If the recording mixes two languages, select the one spoken most.
  • Double-check names, acronyms and technical terms: they're what most often needs a touch-up.
  • To transcribe a video, extract its audio first with Video to Audio.

Frequently asked questions

Which languages does it recognize?

33 in the selector — English, Spanish, French, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic and more — and it can also detect the language automatically.

How do I get subtitles for my video?

Extract the video's audio with Video to Audio, transcribe it here and download the SRT. The file includes the timing of each line and loads straight into your editor or platform.

How accurate is the transcription?

With clear speech and little noise, it's very accurate. Strong accents, background music or people talking over each other increase errors, and long silences can make the model invent words, so review it.

Can I transcribe long recordings?

Yes: the audio is processed in 30-second chunks and you can follow the progress. Ten minutes of audio takes about 4 minutes on a current computer.

Is my audio uploaded anywhere?

No. Transcription happens on your device; only the Whisper model is downloaded, the first time you use it.

How is it different from Text to Speech?

They work in opposite directions: here you turn speech into text, while Text to Speech turns written text into speech.