Click to upload or drop an audio file
MP3, WAV, M4A or OGG — transcribed in your browser, never uploadedAbout this tool
This free tool transcribes an audio file to text — upload a recording and it writes out what was said. It's handy for turning voice memos, interviews, lectures, meetings or podcasts into text you can search, edit and copy.
Beyond a plain transcript, you can switch on timestamps and download the result as SRT or VTT subtitle files — the caption formats used by video editors and YouTube. That makes this a quick, free way to create subtitles and captions from audio, without any separate subtitle software.
The clever part: it runs entirely in your browser. The speech-recognition model (a compact version of OpenAI's Whisper) is downloaded once, about 40 MB, then runs on your own device using WebAssembly. Your audio is never uploaded to a server, so it's completely private — a real advantage over transcription sites that send your recording to the cloud.
Because everything happens locally, transcription speed depends on your device, and longer recordings take longer. After the first run the model is cached, so the tool even works offline. For live dictation from your microphone instead, try the Speech to Text tool.
Features
- On-device transcription — your audio never leaves your device
- Works with MP3, WAV, M4A and OGG files
- Timestamps — toggle time markers to see when each line was spoken
- Four ways to save — plain text, text with timestamps, SRT subtitles or VTT subtitles
- Caption-ready — SRT/VTT files drop straight into video editors or YouTube
- Editable transcript — fix and copy the text before you download
- Works offline after the first run
- Free — no sign-up, no upload, no watermark
How to use it
- Upload an audio file — click the box or drag a file in (MP3, WAV, M4A or OGG)
- Click "Transcribe" — the first run downloads the speech engine (about 40 MB)
- Wait for it to finish — longer recordings take longer to process
- Review and edit the transcript, and tick Show timestamps if you want time markers
- Download as plain text (.txt), or as SRT or VTT subtitles — or copy the text
Frequently asked questions
Is my audio uploaded anywhere?
No. The recording is transcribed in your browser on your own device — it is never sent to a server.
Can I download subtitles (SRT or VTT)?
Yes. After transcribing you can download the result as an SRT or VTT subtitle file, as timestamped text, or as a plain .txt transcript. SRT and VTT include the start and end time of every line.
Can I add timestamps to the transcript?
Yes — tick Show timestamps and each line is prefixed with the time it was spoken. Timestamps are also built into the SRT and VTT downloads.
Can I use this to caption a video?
Yes. Export the transcript as SRT or VTT and load it into your video editor, or upload it to YouTube as a subtitle track.
Why does it download 40 MB the first time?
That's the speech-recognition model. It downloads once and is then cached by your browser, so later runs are quicker and even work offline.
How accurate is it?
It uses a compact Whisper model, which is good for clear speech. Accuracy drops with heavy background noise, strong accents or overlapping voices — and you can always edit the transcript before you download.
What's the difference from Speech to Text?
Speech to Text listens to your microphone live; Audio to Text transcribes an audio file you already have.