The Full Subtitle Translation Workflow

Most people think turning a video into foreign-language subtitles is one job: translate the words. The actual deliverable is a subtitle file with a timeline — miss any step and something reworks itself later. Split the flow into four steps, each with a clear output, finish one before the next, and problems stay visible.

Four steps from transcription to final cut
Source material becomes a source-language SRT by recognition or transcription, terminology is unified, translation controls line length, then export subtitles or burn the final cut

Step one: transcription and proofing — a clean source SRT

The output is a source-language SRT. The classic failure here is "recognition was accurate, but nobody proofed it."

Use SmileSub's subtitle generation to recognize speech from video or audio — source languages include Chinese, English, Japanese and Korean. Proof the result cue by cue: proper nouns, numbers and homophones are where errors cluster. Strip filler words and repetitions while you're at it; they don't belong in subtitles. With noisy audio, recognition will drop or mangle lines — complete the copy before moving on.

A clean source SRT makes every later step cheaper. If the recognition output has overlaps or format problems, run it through the subtitle cleaner first, then translate.

Step two: unify terminology — build a glossary

The output is a term sheet: product names, people, units, brand voice — how everything gets written.

One concept, one rendering, throughout the video. Units and date formats follow the target market, not the source language. Cultural jokes and slang get a decision: localize or cut. The sheet is worth the most in multilingual projects — one glossary feeds every language, instead of each translating alone and drifting apart. For planning multilingual assets, see one video, many language versions.

Step three: translation and line-length control

The output is the target-language SRT. The trap isn't "translating wrong" — it's "the translation doesn't fit."

Translation touches text only; the timeline stays untouched, or lips drift out of sync. Control line length: common guidelines suggest max 18 characters per line for Chinese, 42 for Latin scripts with at most 2 lines. Control reading speed: under 9 chars/second for Chinese, 20 for Latin, each cue holding 1–7 seconds. Bilingual order, once chosen, stays consistent: original on top or translation on top — never switch mid-video.

Produce the target file directly with the SRT translator; for side-by-side original-plus-translation, use the bilingual SRT generator. After export, the subtitle checker batches the line-length and speed audit and lists every offending cue at once.

Step four: delivery — soft track or burn-in

The output is the final video or subtitle file; the platform decides the form.

Platforms with CC support take SRT/VTT — toggleable, and the platform can read the text for recommendations. Vertical social and anything that must be seen burns subtitles into the picture. Web videos take VTT via the SRT to VTT tool; styled delivery takes ASS via the SRT to ASS tool. A uniformly shifted timeline slides with the shift tool — don't hand-edit each cue.

For the burn-vs-soft decision, see burned-in or soft subtitles. For a fuller pre-publish audit, see the video localization checklist.

FAQ

Can transcription and translation be merged into one step?

Technically yes — recognition produces the source SRT, translation produces the target. But skipping the human proof in between carries errors straight into the final cut. Stop at least twice: after proofing, and after the glossary.

Why is line-length control mandatory?

Target-language text is a different length — English usually runs longer than Chinese. Uncontrolled, subtitles overflow the safe area or squeeze into three lines, and readability drops. The checker lists every violating cue in one pass.

Do soft and burned versions both need making?

Depends on the channels. Platforms with CC keep the soft track; social and must-be-seen scenarios add a burn. Both derive from the same SRT — different export forms, no duplicated labor.