Turn speech in any audio file into text — free, no sign-up, nothing leaves your browser.
Both formats pair timestamps with text, but they come from different eras and have slightly different jobs. SRT (SubRip) is the older, simpler format, and it's the one most video editing software and platforms expect when you upload a separate subtitle file. VTT (WebVTT) was designed specifically for the web — it's what HTML5 video players natively support for on-page captions, and it allows a bit more formatting than plain SRT. If you're not sure which one you need, SRT is the safer default for general video editing and upload; VTT is the one to reach for if you're embedding captions directly into a webpage's own video player.
.txt is plain text, no timing — good for reading or pasting elsewhere. .srt and .vtt are subtitle formats with timestamps for each line, the kind video editors and platforms like YouTube expect.
It's genuinely good for clear speech in English, but not perfect — background noise, heavy accents, overlapping speakers, or music can all trip it up. Review the lines before downloading; each one is editable right in the list.
The speech-recognition model has to download to your browser the first time you use it in a session — after that it's cached and much faster.
Yes — every line is directly editable in the list, so you can fix any misheard words or names before downloading the final file.
No. Everything happens locally in your browser, including the transcription itself — your audio is never sent to a server.
We've noticed an ad blocker is active. Cwtch.es relies on ads to stay free — please consider disabling it for this site if you find this tool useful.