Audio to Text
Transcribe speech from an audio file to text, entirely on your device.
The first time you use this tool, a one-time ~110 MB speech-recognition model downloads (bigger than the ~30 MB used by the video/audio conversion tools, since a real speech-recognition model is a lot more data than a codec). It's cached afterwards, and nothing ever leaves your device.
Longer recordings are processed in 30-second chunks, so bigger files just take longer rather than failing.
How it works
- 1Choose an audio or video file.
- 2Pick the spoken language, or leave it on Auto-detect.
- 3Transcribe and copy or download the text.
About this tool
Upload an audio file and get a text transcript of the speech in it — a voice memo, a meeting recording, an interview. A speech-recognition model runs entirely in your browser, so the audio never leaves your device. Works in 99 languages, with automatic language detection.
Is my audio uploaded anywhere?
No, the speech-recognition model runs entirely in your browser — nothing is ever uploaded.
Which languages are supported?
99 languages, either auto-detected or picked manually from the dropdown.
Why is the first use slower than later ones?
A one-time ~110 MB speech-recognition model has to download the first time you use this tool. It’s cached afterwards, so later transcriptions start immediately.
How long can the audio be?
Any length — longer recordings are processed in 30-second chunks, so bigger files just take longer rather than failing.