Three layers of localisation, three different fixes
Localisation happens at three layers, and each has its own moment, its own cost and its own trap. Deciding which layer a change belongs to is most of the work; doing it at the wrong layer is why so many dubbed versions look wrong. As of 2026-09-12.
| Layer | What changes | When | Typical trap |
|---|---|---|---|
| Audio | The spoken performance | After picture lock, or at generation | Expansion makes the new line longer than the mouth |
| Script | Names, address terms, locations, references | At parsing, so it flows downstream | Changing it later means regenerating references |
| Subtitle | Wording, units, equivalents | After picture lock | Cannot touch text visible in the picture |
Inclusion rule. Layers at which a production can be localised. Re-shooting for a market is a fourth option and is outside what generated production usually contemplates. Order. Alphabetical by layer.
1Reading speed is a published limit, not a preference
Distributors publish characters-per-line and characters-per-second limits, and they differ by language and by audience age. A subtitle that fits the frame can still exceed the reading limit, and the limit for younger audiences is meaningfully lower.
A vertical frame tightens this further: the safe area for text is narrow, so the line limit binds before the reading limit does.
Two lines is the working maximum in a vertical frame, and three is not a compromise but a different layout: it covers the part of the picture the composition was built around. Cutting the sentence is cheaper than moving the subtitle.
2Text expansion is why dubbed mouths stop matching
The same sentence is longer in some languages than others, often by a fifth or more. A dubbed line that carries the full meaning no longer fits the mouth it was written for, and the fix is a shorter translation rather than a better dub.
Translating for length as well as for meaning is the single change that removes most of this, and it has to happen before the audio is made rather than after.
The expansion is not uniform across a script either. Short exclamations survive translation at roughly their original length; explanatory lines are where the extra words accumulate, which is a reason to keep exposition out of close-ups in a series planned for several markets.
3Anything visible in the picture cannot be subtitled away
A sign, a note or a screen generated in one language stays in that language. Subtitles cannot reach it, and dubbing cannot either. Fixing it means going back to the script layer, before the shot exists.
That is the strongest argument for keeping legible text out of generated pictures entirely wherever the story does not require it.
Where a sign has to be legible, the cheapest arrangement is to shoot it as its own insert rather than inside a performance shot, so the localised version regenerates one short piece of picture instead of a scene.
4Which layer a change belongs to
Cultural adaptation of names and references belongs at the script layer. Units and idiom belong at the subtitle layer. Performance belongs at the audio layer. A change made at the wrong layer either costs far more than it should or does not hold.
What each vendor publishes about the languages it supports is recorded on the register.
The ordering follows from that: script decisions first, because they constrain the picture; audio second, because it has to fit the picture; subtitles last, because they sit on top of both.
5Where the language counts are recorded
Counts and lists per vendor are kept in the register, each read from the vendor's own domain under one counting rule applied identically to everyone.
- Netflix subtitle timing limits — publishes per-language subtitle timing and line limits
An explainer about method and structure, not a count. No vendor number appears here that is not also in the register. The sourced material is on the register. Related: Three layers, How hreflang works.