A closed caption is one the viewer can switch off, because the text lives in a separate track or file rather than in the picture. YouTube’s CC button, a subtitle track in a video player, and a .srt file next to your video are all closed captions. Here, the SRT and VTT downloads and the soft-subtitle video are closed captions.
An open caption is burned into the picture, so it’s always visible and can’t be turned off. That’s what you want for TikTok, Reels and Shorts, where most people watch with the sound off and platforms don’t reliably show caption files. Our burned-in and animated caption exports are open captions.
Strictly, subtitles assume you can hear the audio and just need the words - often a translation. Captions are for viewers who can’t hear it, so they also describe sound: [door slams], [laughter], and who is speaking. In everyday use, most people say "captions" for the styled word-by-word text on social video and "subtitles" for a plain line at the bottom of the screen, which is how this site uses them too.
Two lines at most, roughly 40 characters per line, broken at natural pauses rather than mid-phrase, and left on screen long enough to read - about a second minimum. Our blocks are built to those rules automatically, splitting on real pauses in the speech rather than on a fixed word count.
Using it
No. There’s nothing to sign up for and nothing to pay. Limits are counted per internet connection rather than per person.
Video (mp4, mkv, mov, webm, avi) or audio (mp3, wav, m4a, aac, flac, ogg), up to 30 minutes and 1 GB. If a file is over either limit we tell you before the upload starts, so you never wait out a long upload to be refused.
Yes, and it’s the thing this tool is best at. Upload the video together with your existing .srt or .vtt - from CapCut, YouTube, anywhere. We keep your exact words, throw the old timings away, and align every word to the audio.
Drop it in as .txt, .docx, .srt or .vtt, or paste it. Transcription is skipped entirely, so nothing can be misheard - no mangled names, drug names, brand names or acronyms - and there’s nothing to correct afterwards.
Twice. You review the transcript before alignment, which is where you fix misheard words. Then in the editor you can change the text of any subtitle block, with the video playing alongside it.
Upload it here, check the words, and download either a subtitle file to load in your editor or the video with subtitles already burned in. If you’d rather do it in your own software, the guides cover DaVinci Resolve, Premiere Pro and iMovie step by step.
Download the SRT here, then in YouTube Studio open your video, go to Subtitles, and upload the file. Uploading a subtitle file beats burning captions in for YouTube specifically: the text becomes searchable, viewers can switch it off, and YouTube can auto-translate it.
Timing and accuracy
They estimate. The transcription model returns text plus an approximate position, and the caption file is built from that guess, which is why captions drift further into a video. We measure instead: forced alignment searches the audio for where each word is actually spoken, down to the phoneme.
Yes, and each part will be perfectly in sync with itself. But the second file’s timings start again at 00:00 - they’re relative to part two, not to your original video - so line each file up against its own clip on the timeline. Split at a pause rather than mid-word, and split without re-encoding so the durations stay exact.
24 languages, plus auto-detect if you don’t choose one. Japanese, Chinese and Thai aren’t supported yet - they’re written without spaces between words, which needs a different approach to breaking subtitles into lines.
Yes. The preview is drawn to the same geometry the renderer uses - same font metrics, same line breaks, same box sizes - so what you style is what gets burned in.
Waiting and limits
No. Short videos are served first, so a quick clip doesn’t queue behind a long lecture. It cuts both ways though: a long video can only be overtaken twice, after which it goes to the front and nothing passes it.
Aligning audio is real computation on real hardware, so jobs run in a queue rather than all at once - five people processing simultaneously would be slower for all five. While you’re waiting, the screen shows exactly how many videos are ahead of you.
Less time than the video runs for. An 18-minute video came back in under three minutes in our own testing, and a 5-minute one is typically done in about a minute. Bringing your own script skips the transcription step, so it is quicker still. If other jobs are ahead of yours, the waiting screen shows how many.
It’s what keeps the service free and available to everyone. The window rolls rather than resetting at midnight, so a slot opens up four hours after each job you started.
Files and privacy
It’s processed on our servers and deleted when you close the job page, along with everything made from it. Leave the page open and your files stay while you work, so take the time you need. It is never used to train anything, never shared, and never shown to anyone else. There are no accounts, so there’s no profile to attach it to.
SRT and VTT subtitle files, your video with a soft subtitle track it can be turned off, your video with subtitles burned in, or a version with animated word-by-word captions in one of six styles.
Yes. SRT comes in two variants because editors disagree about line endings - Resolve wants CRLF, AutoSubs wants LF - and both are one click apart on the download menu. The guides walk through each editor.
It depends which kind you have. A closed caption - a separate SRT, or a soft subtitle track - is just switched off or deleted, and the video underneath is untouched. Burned-in captions are part of the picture and can’t be removed; you need the original video back. Worth knowing before you burn anything in.
Yes. It’s your video and your words; we claim nothing over them and ask for no credit.