
Captions, subtitles, SDH, and forced narratives aren't the same deliverable
Before choosing a file format, make sure everyone is talking about the same kind of timed text. People use the words loosely in post, but delivery specs often use them very specifically.| Timed text type | Primary viewer need | Usually includes | Common delivery note |
|---|---|---|---|
| Closed captions | Accessibility for deaf or hard of hearing viewers | Dialogue, speaker IDs when needed, meaningful music, sound effects, off-screen speech | Often subject to strict platform or broadcast caption rules |
| SDH subtitles | Accessibility-style subtitles, often in subtitle workflows | Dialogue plus relevant non-speech audio and speaker context | May be requested instead of traditional closed captions by streamers |
| Full subtitles | Translation or transcription for viewers who can hear the audio | Spoken dialogue, sometimes translated on-screen text | Not the same as SDH unless non-speech accessibility cues are included |
| Forced narratives | Limited translation for foreign dialogue or necessary on-screen text | Only the lines or text the primary-language viewer needs | Should not be delivered as a replacement for full subtitles or captions |
| Dub subtitles | Subtitle support for a dubbed audio version | Text matched to the dub script or dub timing, depending on spec | Confirm whether the platform wants dub-matched text or original-language subtitle translation |
- English closed captions or English SDH for accessibility
- Full subtitles for translated languages
- Forced narrative subtitles for foreign dialogue or important on-screen text
- Dub subtitles, if required by the platform or localization workflow
- Textless elements or reference notes for graphics that are handled outside the subtitle file
Pick the format from the destination
Most video editing tools can export several caption formats. That doesn't mean all of them are acceptable for delivery. A format that works fine for review on YouTube may be wrong for a studio, broadcaster, streamer, or aggregator. Here are the common formats you'll run into in delivery conversations:| Format | Extension | Common use | Main limitation |
|---|---|---|---|
| SRT | .srt | Web platforms, screeners, simple sidecar subtitles | Minimal styling, limited positioning, no rich metadata |
| SCC | .scc | North American broadcast-style closed caption workflows, CEA-608 heritage | Legacy constraints, limited character set and formatting |
| STL | .stl | EBU subtitle delivery, common in some broadcast and international workflows | Region/spec dependent, can be easy to misconfigure |
| TTML / IMSC / SMPTE-TT | .ttml, .xml, sometimes .dfxp | Streamers, IMF-adjacent workflows, modern platform delivery | More complex, strict validation, spec-specific profiles |

Sidecar delivery is usually safer than embedding
For professional delivery, sidecar timed text files are usually the cleanest path. A sidecar file travels next to the video master as a separate asset, and it can be replaced, versioned, language-tagged, validated, and re-delivered without re-encoding picture. Embedded captions live inside the video file. That can be useful for certain broadcast masters or archive workflows, but it's risky if the destination doesn't want embedded tracks. Some services explicitly reject video submissions that contain embedded closed caption tracks and require captions as separate SCC or TTML files. Others may accept embedded captions only for specific containers or distribution paths. Burned-in subtitles are different again. They're rendered into the picture and can't be turned off. Use them only when the spec or creative intent calls for open captions, burned-in foreign dialogue, festival screeners, or review copies. Burned-in text is almost never a substitute for accessibility captions because the platform can't expose it as a selectable caption track. The delivery choice usually falls into these buckets:- Sidecar captions or subtitles for most streamers, platforms, aggregators, and localization workflows
- Embedded captions when a spec explicitly requests them in the video container
- Burned-in text for open-caption versions, festival copies, or creative on-screen translation
- Both sidecar and burned-in reference only when the recipient asks for both and understands which is authoritative

Timing has to match the delivered master
Caption sync issues often come from using the wrong reference file. A vendor may time captions to an H.264 review export that starts at 00:00:00:00, while the final mezzanine starts at 01:00:00:00 and includes two seconds of black at head. Or editorial may trim a logo after captions were approved. Or a distributor may require bars, tone, slate, or act breaks that weren't in the caption reference. Your team should conform every timed text asset to the length and timing of the accompanying video. Some platform specs allow a small tolerance, such as subtitles conforming within half a second, but you shouldn't use that as working slack. Half a second late is very visible in dialogue, and a few frames can matter on fast exchanges, comedy timing, or music performance. Your team should verify timing in the context where the file will be delivered:- Confirm the video file used for caption timing is the same version as the delivery master
- Confirm runtime, start timecode, and frame rate
- Confirm captions don't start before first frame of program or extend past end of file
- Confirm any slate, logo, recap, teaser, credits, or post-credit scene is handled correctly
- Confirm mastering didn't shift act breaks, commercial blacks, and texted sections

Text rules are part of delivery
Caption and subtitle text has to be readable under real playback conditions. That means line length, duration, line breaks, speaker labels, punctuation, and placement all affect whether the asset passes QC and whether viewers can use it. Different platforms publish different style rules, but most specs are trying to control the same things:- Maximum characters per line
- Maximum lines per subtitle or caption event
- Minimum and maximum duration per event
- Reading speed, often measured as characters per second
- Line break rules for grammar and readability
- Treatment of speaker IDs, off-screen dialogue, and overlapping speakers
- Treatment of music, lyrics, and sound effects
- Placement to avoid lower thirds, burned-in text, credits, and other plot-relevant graphics
- Language-specific punctuation, casing, and character support
QC captions against picture, audio, and the delivery spec
Caption QC is a three-way comparison between the timed text file, the final picture/audio, and the delivery requirements. A caption file can be beautifully written and still be the wrong asset type, wrong language, wrong timebase, or wrong profile.
- Technical validity: file opens, parses, validates, and matches the requested extension and format
- Asset identity: language, territory, version, episode number, cut name, and file naming match the delivery request
- Timing: captions appear and disappear in sync with speech and meaningful audio
- Completeness: all required dialogue, SDH information, forced narrative text, credits, and plot-relevant on-screen text are covered
- Readability: line length, reading speed, line breaks, and event duration feel usable
- Placement: captions avoid burned-in text, lower thirds, credits, and important visual information where the format supports positioning
- Text accuracy: spelling, names, lyrics, terminology, censored words, and punctuation match the approved program and spec
- Playback behavior: the file displays correctly in a player or validation tool that's relevant to the destination
Version control prevents most late-stage caption pain
Captions usually arrive late because they depend on locked picture, final audio, approved titles, and sometimes legal or network notes. That makes version control critical. If your team is casual about filenames and references, captions become a guessing game. Use filenames that make the asset identity obvious. Include project, episode or reel, language, timed text type, version, frame rate if useful, and date or revision. Keep the naming consistent with the delivery spec when one exists. For example, your team should be able to tell the difference between:- English SDH captions for the final domestic master
- English dialogue subtitles for a festival screener
- Spanish full subtitles for the international texted master
- English forced narrative subtitles for foreign dialogue only
- Revised captions conformed after a picture trim
A clean handoff makes caption delivery boring
The best caption deliveries are uneventful. The post supervisor knows the required formats early, and editorial and finishing know which master is authoritative. The caption vendor receives a stable reference with accurate timing information, and QC reviews the exported sidecar files against the final picture. Nobody is trying to turn a web SRT into a platform-specific TTML file at midnight. When you set up the workflow, define these details as soon as you have delivery specs:- Required timed text types, such as CC, SDH, full subtitles, and forced narratives
- Required languages and territories
- Required file formats and profiles
- Whether delivery is sidecar, embedded, burned-in, or some combination
- Master video runtime, frame rate, and start timecode
- Naming rules and version identifiers
- Who approves text accuracy, accessibility treatment, and final sync
- Which file is the timing source of truth
FAQ
Closed captions and SDH are accessibility tracks for viewers who are deaf or hard of hearing, so they include dialogue plus meaningful non-speech audio such as speaker IDs, sound effects, music cues, and off-screen speech. Subtitles usually provide dialogue translation or transcription for viewers who can hear the audio. Forced narratives are limited subtitle events for foreign-language dialogue, signs, or on-screen text that must be understood when watching in the primary language.
Use the method required by the delivery spec. For many professional deliveries, sidecar files are preferred because they can be versioned, replaced, language-tagged, validated, and redelivered without re-encoding the picture. Embedded captions are appropriate only when the destination explicitly requests them in a supported container and caption standard. Burned-in subtitles should be used only for open-caption versions, screeners, festival copies, or creative translation that's meant to be permanently visible.
SRT is simple and widely supported, but it carries limited styling, positioning, metadata, and validation information. A review platform may display it acceptably, while a broadcaster, streamer, or aggregator may require SCC, STL, TTML, IMSC, or a specific platform profile. SRT can also hide readability problems such as long lines, poor wrapping, missing speaker treatment, or missing SDH information.
Load the exported caption or subtitle file against the exact video master being delivered, not just the NLE timeline or a review proxy. Confirm runtime, frame rate, start timecode, drop-frame or non-drop-frame behavior, head and tail timing, act breaks, logos, credits, and any late trims. Spot-check the beginning, middle, dense dialogue scenes, music sections, credits, tail, and every area affected by picture changes.
Common rejection reasons include the wrong format or profile, wrong language or asset type, mismatch to the delivered master, incorrect frame rate or timecode basis, captions extending past the program end, missing SDH cues, missing forced narratives, excessive line length, poor reading speed, invalid characters, bad positioning, and filenames or metadata that don't match the delivery request.
Caption notes should be tied to the picture frame where the issue appears, not described only in email or a spreadsheet. Aspect supports frame-accurate comments, so a sync issue, typo, bad line break, or missing SDH cue can be reviewed at the exact point in playback.





