AreaAudit

Which languages each generator says it supports, and where

Speech synthesis: building a voice, one language at a time

Speech synthesis turns text into audio. Building it for a language needs a voice, licensed and recorded, which is why a list of languages a tool can speak is structurally shorter than a list of the languages it can hear. As of 2026-09-22.

Why the output inventory is structurally the shorter oneRecognition trains on recordings of people talking, which exist in quantity in many languages. Synthesis needs a voice recorded for the purpose with a contract behind it, per language and often per variety, so the cost per language is far higher.Producing speechRecognising speechTraining materialA voice recorded for the purposeRecordings of people talkingLicensingPer voice, and negotiatedRarely per speakerCost per languageHighLowerTypical inventorySmallerLargerA jump in headline coverage most likely extended the cheaper side
Fig. 1 Every vendor selling this has the gap, whether or not its pages show it. One vendor here publishes both sides and the difference is more than forty languages.
Why synthesis inventories lag recognition inventories. Recorded 2026-09-22.
RequirementSpeech synthesisSpeech recognition
Training materialA voice recorded for the purposeRecordings of people talking
LicensingPer voice, and negotiatedRarely per speaker
Cost of adding a languageHighLower
Typical inventory sizeSmallerLarger

Inclusion rule. The two directions of a dubbing pipeline, compared on what adding a language to each one requires. Order. Fixed order, the requirement that differs most first.

1A voice has to be built and paid for

Recognising speech can be trained from material that already exists in quantity. Producing it needs a voice recorded for that purpose, with a contract behind it, per language and often per variety.

So the asymmetry between the two sides of a dubbing product is structural rather than a gap somebody forgot to close. Every vendor selling dubbing has it, whether or not its pages show it.

2One vendor here publishes both sides

A single vendor in this register prints two lists, one of the languages it can detect in a source file and a shorter one of the languages it can dub into. The difference is more than forty languages.

That vendor is not unusual in having the gap; it is unusual in showing it. Every other spoken-language figure here is one-sided, and a reader cannot tell which side of the pipeline it describes.

3Which side a figure describes changes what it is worth

A figure describing recognition tells a buyer what material they can bring. A figure describing synthesis tells them which markets they can reach. Only the second one has revenue attached to it.

When a vendor publishes one number and no direction, this register records the wording rather than guessing which side it belongs to. The wording is usually the only clue, and sometimes there is none.

A definition, not a measurement: no vendor figure appears here that is not also in the register. The one two-sided list in this register is recorded on that vendor cell.

Nearby terms: Voice inventory, Speech recognition. The whole vocabulary is at terms; nothing on this page is a claim about a named product.