How to turn speech into text
- Open a file or record. Pick an audio or video file (MP3, M4A, WAV, OGG, WebM, MP4, MOV) or click Record from microphone and talk. Either way the sound stays in this tab.
- Choose the language and transcribe. Pick the language spoken in the audio. Tick Label speakers if more than one person talks. The first time, the speech model downloads once; after that it starts from your browser’s cache.
- Copy or download. Fix any misheard words in the box, then copy the text, save it as a .txt file, or save .srt subtitles with timings for a video.
Good to know
- Three model sizes. Quick (41 MB) suits clear English. Better (77 MB) is more accurate for English and major European languages. Best (249 MB) is slower but the one to use for Hindi and other Indian languages; the smaller ones often write them in English letters or translate them.
- Speaker labels work best on short clips. They are made for up to 3 people and recordings of a few minutes. When two people talk over each other, the sentence goes to whoever spoke more of it.
- Long recordings work in parts. The audio is split at natural pauses into pieces of under 30 seconds, and the text appears as each piece is done. On a recent laptop, Better takes about half as long as the audio itself and Best a little longer than the audio; older computers and phones are slower.
- Clear audio, better text. One speaker close to the microphone gives the best result. Music, crosstalk and echo lower the accuracy.
- Timestamps. Tick Show timestamps to see when each line starts. The .srt download always has timings, whatever you choose.
Your audio stays on your device
Most transcription sites upload your recording to their servers, and even your browser’s built-in dictation can send your voice to the cloud. Here, the Whisper speech model runs inside this browser tab. Your file or recording is decoded, transcribed and saved on your own device and is never uploaded.
This tool downloads something on first use: your browser fetches the open-source Whisper model from Hugging Face and the ONNX runtime from jsDelivr. Those are downloads to you, not uploads of your audio. The microphone is only on while the button says Stop recording. You can check: open your browser’s developer tools, watch the Network tab, and transcribe something. How InTheTab works
You’ll also see a few small requests to Google Analytics. That’s our visit counter, and it never receives your files. Privacy policy
Questions
Is my audio uploaded to a server?
No. The Whisper speech model runs in your browser, so the audio never leaves your device. The only network traffic is the one-time download of the model from Hugging Face and its runtime from jsDelivr. We never get your files or text: the site has no accounts, and its only analytics, Google Analytics, counts visits without ever seeing what you add here.
Why does the first use take longer?
Your browser has to download the speech model once: about 41 MB for Quick, 77 MB for Better or 249 MB for Best. After that it loads from your browser’s cache, so later uses start much faster.
Can it transcribe Hindi?
Yes. It understands Hindi, English and more than 90 other languages, and you choose which one is spoken. For Hindi and other Indian languages, choose the Best model: the smaller ones often write Hindi in English letters or translate it.
How do I make subtitles for a video?
Open the video file, click Transcribe, then click Download .srt. Most video players and editors, including VLC and YouTube Studio, accept SRT subtitle files.