AreaAudit

Which languages each generator says it supports, and where

Speech recognition: the input side of a dubbing product

Speech recognition turns recorded audio into text. It is the first step of dubbing a finished file, its language list is normally the longer of the two, and it constrains what a buyer can bring rather than what they can deliver. As of 2026-09-22.

Which side of a pipeline each list constrainsThe input list limits which recordings can be processed, so it decides which catalogues a buyer can bring. The output list limits which languages a result can be in, so it decides which markets can be reached. Only the second has revenue attached.Input listLanguages the system can readLimits the materialWhich catalogues can be processedOutput listLanguages it can deliverLimits the marketsWhat the buyer is purchasingOne operation, and only one of the two lists is the offerWhat each side decides
Fig. 1 Where a vendor publishes one figure and no direction, a reader has a number whose side and whose unit are both unstated.
What each side of a dubbing pipeline constrains. Recorded 2026-09-22.
SideConstrainsCommercial meaning
RecognitionWhich recordings can be used as inputWhich catalogues can be processed
SynthesisWhich languages the result can be inWhich markets can be reached
Both togetherThe whole operationThe product, as sold

Inclusion rule. The two sides of a dubbing operation and what each one limits. Both lists are needed before either figure means anything. Order. Input side first.

1The longer list is the less useful one

A vendor that can recognise seventy languages and speak twenty-nine sells into twenty-nine markets. Quoting the seventy would describe its transcription and appear to describe its dubbing.

Where both sides are published, this register records the smaller figure as the entry and prints the larger beside it. That ordering is the whole reason to publish both.

2Recognition quality is not a list either

A language appearing on a recognition list says the system will attempt it. Accented speech, overlapping dialogue, background music and archive audio all change the result and none of them appears in any figure.

For a catalogue being localised, that matters as much as the list does. Nothing in this register addresses it, because nothing published by any vendor here does.

3It is the cheaper half to extend

Adding a language to recognition needs recordings, which exist in quantity for a large number of languages. Adding one to synthesis needs a voice built and licensed.

So a vendor announcing a large jump in language coverage has most likely extended the input side. A reader can test that reading by looking for which of the two lists moved, if the vendor publishes both.

A definition, not a measurement: no vendor figure appears here that is not also in the register. The gap between the two sides is recorded on the one vendor that prints both.

Nearby terms: Speech synthesis, Source language. The whole vocabulary is at terms; nothing on this page is a claim about a named product.