Speech to Text from an Audio File - Free Toolbox
Get speech to text from an audio file with no software to install. Lectures, interviews and voice memos become timed text, subtitles or a Word file, free.
- Free to use, no login
- Runs locally in your browser
- Clean output, no stamp
Using this tool means you accept our Terms and Privacy Policy.
Big files need a moment. Please leave this tab open until it finishes.
Your file is ready
Something went wrong
Turn a recording into written text
Drop in a voice memo, a lecture, a podcast episode or a video clip. Several files can go in the same batch.
Choose between the quick and the precise speech model, then name the language or leave detection on.
Read along, fix anything that sounds off, and save the result as plain text, subtitles or a DOCX file.
What this toolbox does with your recordings
Phone memos and video clips welcome
M4A memos, MP3 podcasts, WAV dictation, FLAC and OGG files all work, and so does the sound inside MP4, MOV or MKV footage. Rare formats like WMA and AMR get converted locally first.
Pick speed or precision
The quick model (Whisper tiny) suits clean speech. The precise one (Whisper base) copes better with accents, noise and smaller languages, at around half the pace.
Ready for captions or reports
Export SRT or VTT caption files with two-line cues and a line width of 32, 42 or 50 characters, or a DOCX or TXT document, with or without time codes.
A checker built in
Every line sits next to its time code. Press the time to hear that moment again, retype a wrong word, or swap a misspelled name in every line with one Replace all.
Speech to text from an audio file: real jobs it handles
Typing up a recording by hand takes several times its length. Letting a speech model draft the text and then correcting it is much quicker, and here the model works inside your browser, so the file stays on your computer or phone.
- Students: record a lecture, get a draft of the notes, and search it for the topic the teacher said will be on the exam.
- Journalists and researchers: draft interview transcripts with time codes, so each quote can be checked against the audio before publishing.
- HR and office teams: turn recorded meetings or candidate interviews into a DOCX draft without sending sensitive talk to an outside service.
- Freelancers and creators: produce SRT captions for a podcast clip or a short video, ready to load into an editor or a video platform.
- Teachers: caption recorded lessons in VTT for web players, or hand students a printable text version.
A workflow that saves time
Start with the quick model and press Stop after a minute or two: the text so far is kept, so you can judge it early. If names and terms come out wrong, switch to the precise model for the full file. After it finishes, use Replace all for any name the model keeps misspelling, then listen only to the lines that look odd instead of replaying everything.
Where it struggles
People talking over each other, loud music and poor phone lines lower the quality, and there are no speaker labels, so add names yourself if you need them. Recordings of several hours on a phone may run out of memory; cut them into parts of an hour or less. Processing on your own device is also slower than a paid cloud service, and the screen is kept awake while it works.
Other media and QR tools in the box
Questions people ask about this task
Can I convert a voice memo from my phone into text?
Yes. Phone memos are usually M4A, which is supported directly. Open this page on the phone or send the memo to a computer, which is faster, and add it here.
Do I need to install any software or create an account?
No. Everything happens in the browser tab. The first run fetches the speech model (about 41 MB quick or 77 MB precise), and your browser keeps it for next time.
Can I transcribe several recordings in one go?
Yes. Select multiple files and they are processed one after another. Each gets its own transcript to review, and the finished files can be saved one by one or all together.
Can it translate a recording into English?
It can. Set Result to Translate to English and spoken Spanish, German, Romanian or any other supported language comes out as English text. Treat it as a rough draft, since the models are small.
What if the tool guesses the wrong language?
It shows how confident the guess was and warns you when it is unsure. Use Change settings, pick the correct language from the list of 99, and run it again.
Why is part of my recording missing from the text?
Either the run was stopped early, in which case a note says where it ended, or a long quiet stretch was skipped. Turning off Skip silent parts makes the tool process every second.