Videos · Compared
Subtitles vs captions: what is the difference?
Captions transcribe a video's audio in the same language, including important sounds and speaker changes, for people who cannot hear it. Subtitles translate the speech into another language for people who can hear but do not understand it. SDH, subtitles for the deaf and hard of hearing, combine the two: subtitle styling with sounds and speakers included.
The three look identical on screen, a line or two of text near the bottom of the frame, so they get labelled interchangeably. They are written for different viewers, and a file made for one audience quietly fails the other.
Nuwan Madhusanka · Co-founder
5 min read · Published
| Captions | Subtitles | SDH | |
|---|---|---|---|
| Language | Same as the audio | Translated into another language | Usually the same as the audio, sometimes translated |
| Written for | Viewers who cannot hear the audio | Viewers who hear but do not understand the language | Deaf and hard of hearing viewers |
| Sound effects and music | Included when needed to follow the video | Left out | Included |
| Speaker identification | Included when the speaker is not obvious | Rarely | Included |
| Typical home | Web accessibility and North American broadcast | Foreign language films and series | Streaming and disc releases |
| UK usage | Usually called subtitles | Subtitles | Subtitles for the deaf and hard of hearing |
Three jobs for the same strip of text
The useful way to separate the terms is to ask who the text is for. A caption stands in for the soundtrack, so it has to carry everything a hearing viewer would take from the audio: the words, who says them, and sounds that change the meaning, such as a phone ringing or music turning tense. A subtitle stands in for language knowledge, so it assumes the viewer hears the sounds and only needs the dialogue in a language they read. The W3C puts it simply: captions are for audio in the same language, subtitles for spoken audio translated into another language.
What goes into each
Captions include dialogue, speaker labels where the speaker is off screen or unclear, and short sound descriptions in brackets, like applause or a door closing. Music is noted when it matters, for example a song that carries a lyric the story depends on. Subtitles usually contain dialogue only, and translators often condense lines to fit a comfortable reading pace, because translated sentences can run longer than the original. SDH follows subtitle conventions for timing and layout but adds the speaker labels and sound descriptions captions carry, so that a track made for a film release also serves viewers who cannot hear it.
The words change between countries
Much of the confusion is regional. In North America, captions means same language text for accessibility, and subtitles means translation. In the UK and much of Europe, subtitles is used for both, and the BBC publishes its accessibility guidance as subtitle guidelines, covering timing, colour and sound labels for deaf and hard of hearing viewers. SDH grew up as a label on streaming services and disc releases, where a subtitle file needed to signal that it included sound information. When a brief, a contract or a platform setting uses one of these words, check which meaning is intended before delivering a file.
When it matters
It matters whenever a video carries information someone must receive. The W3C accessibility guideline for prerecorded video asks for captions on synchronised audio, which a translated subtitle track does not satisfy on its own, because it leaves out the sounds. It matters for reach too: a video sold into another market needs subtitles or a new voiceover, and captions in the original language do nothing for that audience. And it matters in feeds, where many hearing viewers watch with the sound off and read captions simply because audio is not an option in that moment.
Common mistakes
Labelling a translated subtitle file as captions is the most common, and it leaves deaf viewers without sound cues. Publishing machine transcribed captions without a check is the next, since product names, numbers and accents are exactly where recognition goes wrong. Summarising instead of transcribing turns captions into a paraphrase that no longer matches the lips on screen. Mixing languages within one track, such as translating some lines and leaving others, confuses both audiences. And leaving out a crucial sound, such as a warning tone in a safety video, removes the one piece of information that viewer most needed.
Where it shows up in the product
The video builder here makes captions in the strict sense: they are generated from each scene's narration text, so they are always in the language that scene is narrated in, and they carry speech only, with no sound descriptions or speaker labels. Language can be set per scene, so captions follow if narration switches language partway. There is no separate subtitle track and no way to caption in a language different from the voice. To reach another market, build a version narrated in that language, as the Spanish restaurant example does, and the captions arrive in Spanish with it. The captions are burned into the picture rather than supplied as a file.
Questions people ask
Do captions need to describe music?
Only when it matters to understanding. A brief note such as upbeat music is enough where a mood shift carries meaning, and lyrics should be captioned when the words are part of the message. Background music under narration that changes nothing about the content usually does not need a label, since describing every bar would crowd out the dialogue.
Are automatic captions good enough?
They are a starting point rather than a finished track. Speech recognition handles clear, single speaker narration reasonably well but struggles with brand names, numbers, technical terms, accents and overlapping voices. Read the whole track against the audio before publishing, and correct every name, figure and date, because those are the errors that mislead viewers rather than merely looking untidy.
What is the difference between captions and a transcript?
Captions are timed to appear with the speech while the video plays. A transcript is the full text presented separately, usually on the page beside the video, with no timing. Transcripts help people who prefer reading, who use a screen reader, or who want to search or skim content quickly. Many accessible pages provide both, since each serves viewers the other does not.
Can one video carry several subtitle languages?
Yes, on platforms that accept caption files. Each language is uploaded as its own track and the viewer picks one from the player menu, so a single video serves several markets. That only works with closed tracks. Text burned into the picture is fixed to one language, so every additional language burned in means a separate video file.
What are forced subtitles?
Forced subtitles appear only for parts of a video that the main audience would not otherwise understand, such as a short exchange in another language or a sign written in a foreign script. They display even when a viewer has turned subtitles off. They are not an accessibility track, since they cover only those moments rather than all of the dialogue.
Make one with videos
The button opens the generator with this use case already described. Change the wording to match your own.
Create a video with OneCraftRelated questions
- Burned in captions vs closed captions: what is the difference?Burned in captions vs closed captions: text in the picture versus a track viewers can switch off. How each works, where each fits, and the mistakes to avoid.
- What is an SRT file?What is an SRT file? The plain text SubRip caption format explained cue by cue, how it differs from WebVTT, and the mistakes that break captions on upload.
- What is text to speech (TTS)?What is text to speech? How TTS turns writing into spoken audio, how concatenative, parametric, neural and generative voices differ, and how to write for them.
Step by step in the builder: Add voiceover, captions and music to a video.
Written and checked by the OneCraft team. Last checked .