The terms get used interchangeably. They aren't the same. A Closed Caption is a transcription of all audible content - dialogue, music, sound effects, plus speaker identification - designed for viewers who can't hear the audio. A Subtitle is a transcription of just the dialogue, designed for viewers who can hear the audio but don't speak the language.
Most short-form video creators want subtitles. Most accessibility regulations require captions. Knowing which is which matters more than getting either one technically perfect.
The technical difference
A working definition of each:
Closed Caption (CC) - burned-in or selectable text that includes:
- Dialogue, with speaker IDs when ambiguous.
- Music cues (e.g., [soft piano music]).
- Sound effects (e.g., [door slams]).
- Non-speech vocalizations (e.g., [laughs], [sighs]).
- Tone or volume cues when relevant (e.g., [whispers], [over the phone]).
- Dialogue only.
- Speaker IDs only when necessary for the dialogue to make sense.
- Translations when the audience speaks a different language than the source.
When you need captions, when you need subtitles
Three cases that determine the answer:
- Distribution on US TV, Netflix, broadcast platforms. Captions are legally required (FCC, ADA). Subtitles alone don't qualify.
- Distribution on YouTube, TikTok, Instagram for an audience speaking the source language. Subtitles are the working format - viewers in noisy environments who can hear but choose not to.
- Distribution to non-source-language audiences. Subtitles in the destination language, translated.
The file format question
Two formats dominate:
SRT (SubRip Text) - plain text, time-coded, no styling. The most-supported format across platforms. Looks like:
1
00:00:00,000 --> 00:00:02,500
Most explainers fail at the hook.
2
00:00:02,500 --> 00:00:04,200
The hook is three seconds long.
VTT (Web Video Text Tracks) - similar to SRT but supports basic styling, positioning, and cue identifiers. Required for HTML5 video; usable on most platforms.
For uploading to YouTube, TikTok, Instagram, X: SRT is universal. For embedding video on your own site with HTML5: VTT. The SRT Converter handles conversions between SRT, VTT, and the various proprietary caption formats (Final Cut, Premiere, DaVinci).
Burned-in versus selectable
Two ways to deliver captions or subtitles:
Burned-in (open). The text is baked into the video pixels - can't be turned off. Pros: works everywhere, no upload complications, controls styling. Cons: can't be turned off (problem for users who don't need them), can't be translated post-upload.
Selectable (closed). Caption file uploaded alongside the video; viewer toggles on/off. Pros: respects user preference, multiple languages possible. Cons: requires platform support, styling is platform-controlled.
For short-form social video, burned-in is the default. Mobile viewers in public environments tend to watch with sound off - captions need to be there from frame one, not opt-in.
For long-form YouTube content, selectable is the default. Long-form viewers tend to watch with sound on; captions are a fallback. Selectable captions also enable YouTube's auto-translate.
For accessibility-mandated content, both. Burned-in for default visibility, plus a separate caption file for screen readers and assistive technology.
Typography for burned-in captions
The rules from the lower-third guide apply double for captions, because captions are read more attentively:
- Sans-serif, bold, single weight.
- Two lines maximum per cue.
- High-contrast scrim or outline.
- 80% screen height position (bottom of the lower third).
- Maximum 32 characters per line. Anything longer scrolls badly on a phone and reads slowly on a desktop. The Subtitle Translator flags lines that exceed this limit when translating.
- Cue duration: 1.5–7 seconds. Shorter and the viewer can't read it; longer and the brain forgets the first half by the time the second half lands.
- No more than 25 lines per minute of video. Above that, the caption rate is faster than reading speed and viewers fall behind.
Translation versus transcription
Translating subtitles is a different job than transcribing them. Three tactical notes:
- Don't auto-translate without a human pass. Machine translation handles the literal meaning, but misses idiom, tone, and cultural reference. A 60-second video's subtitle file is small enough that a human review is fast and worth doing.
- Account for character expansion. Translations from English to German typically expand by 20-30%; English to Spanish by 25%; English to French by 30%. Lines that fit the 32-character limit in English may break that limit in translation.
- Reading speed varies by language. Languages with shorter words (English, Chinese) can support faster cue rates than languages with longer words (German, Russian). Hand-edit cue durations after translation if you notice viewers struggling.
Accessibility minimums
Three things that move a caption file from technically present to actually useful:
- Speaker identification when there are multiple speakers. "[Sarah]" and "[interviewer]" in front of each line.
- Sound effect annotation when the effect matters. Don't transcribe ambient noise; do transcribe [phone rings] or [door closes].
- Music cues for songs with lyrics. "♪ [song title by artist] ♪" or "♪ lyrics ♪" on a separate line.
Burning captions in cleanly
Two technical pitfalls when burning captions into the video:
- Render at the target resolution, not the upload resolution. If you're uploading 1080p, render the captions at 1080p. Scaling captions down from 4K to 1080p softens the edges and reduces legibility.
- Anti-alias the text against a scrim, not the raw image. Captions on top of a busy background, even with outline, struggle. A 30%-opacity dark scrim across the bottom 30% of the frame, with the captions on top, is the cleanest combination.
A working pipeline
For a typical short-form video:
- Generate auto-captions in your NLE or via the platform's tools.
- Hand-edit for accuracy, capitalization, and line breaks.
- Add speaker IDs and sound cues if the destination is accessibility-mandated.
- Burn captions into the export for social platforms.
- Upload SRT separately for YouTube/long-form, where selectable captions are the default.
- Translate via Subtitle Translator for destination-language audiences, then human-review.
Bring it to VideoCue
The SRT Converter handles SRT/VTT/proprietary-format conversions, and the Subtitle Translator handles destination-language translation with character-expansion awareness. The Filmwiki entries on Closed Caption and Subtitle cover the regulatory and accessibility distinctions in more depth - useful when you're shipping to a platform with mandated requirements.