AreaAudit

Which languages each generator says it supports, and where

The vocabulary a language list is written in

These pages define the words a language claim is made of. None of them contains a figure about a named product, because a definition that depends on one vendor is not a definition. As of 2026-09-12.

Why one word has to be split into several termsA single claim of language support can describe a translated menu, a generated script or a synthetic voice. The sentence stays true of one and false of the others, so the vocabulary has to be split before any cell can be filled.One vendor sentenceA claim of support,with no statement ofwhich system itdescribes.Which system?Split by termInterface, generateddocument, or spokenvoice, each with itsown definition.Is a listactually named?A cell that can becheckedThe claim aspublished, filedunder one column,with its readingdate.
Fig. 1 The terms section exists so a register cell can stay short: the distinction it relies on is defined once, here.

Almost every argument about language coverage is a vocabulary problem. One word covers a translated menu, a generated script and a synthetic voice, and a sentence using it can be true of one and false of the others without changing a character.

So each term is given its own page with the same three parts: what it means, what it pins down, and what readers take it to promise that it does not. Where a term has a narrow technical sense and a loose marketing sense, both are written down and the gap between them is the point.

These pages are the layer the register leans on. A cell can stay short because the distinction it depends on is defined once here instead of being restated in a parenthesis on every row.

1The vocabulary

  • Language tag — The complete code a page or a voice is labelled with, which parts of it are required, and how much a reader can infer from the parts that are missing.
  • Language subtag — The two or three letters naming a language, why some languages have two codes, and what a comparison of bare codes can and cannot conclude.
  • Region subtag — The country or region appended to a language code, what it pins down for a reader, and why most of the vendors in this register leave it off.
  • Script subtag — The part of a code that names a writing system, the few languages where it is necessary, and why it is not interchangeable with a region.
  • hreflang attribute — The markup that tells a search engine which address serves which language, why this register counts it, and the four questions it cannot answer.
  • x-default — The alternate entry that names no language, what it is for, and why counting it would measure configuration quality instead of language coverage.
  • lang attribute — The attribute in which a document states its own language, how it differs from an alternate declaration, and what it is worth as evidence.
  • Locale — The bundle of language, region and formatting conventions a product switches between, and why a locale count is not a language count.
  • Language cluster — A group of addresses declaring each other as language versions of one page, why vendors maintain them, and what a broken cluster looks like.
  • Language switcher — The menu a visitor uses to change a site language, why it is not what this register counts, and what it can still tell a reader.
  • Machine-swapped page — A page put through automatic translation and published unreviewed, how to recognise one, and why a declared alternate cannot rule it out.
  • Language subdomain — Serving each language from its own subdomain, why the counting rule discards those alternates, and what the discarded figure still tells a reader.
  • Country-code domain — A separate national domain for a market, why an alternate pointing at one is not counted here, and what cannot be read from the markup.
  • Named list — A published set of language names, why it is worth more than any count, and the four forms it takes among the vendors read here.
  • Bare count — A language figure published without the set it counts, why the size of the number is irrelevant, and why the register still records it.
  • Plus figure — The habit of writing a language count with a plus after it, what that does to the claim, and why it removes the pressure to keep a page current.
  • Printed total — A list carrying its own count, why it is the most useful shape of disclosure in this register, and how rarely any vendor bothers with it.
  • Per-job cap — A limit on how many languages a single job can produce, why it affects a schedule more than an inventory figure does, and how rarely it is published.
  • Empty cell — What a blank in this register asserts, what it deliberately does not, and why the distinction changes how a whole column should be read.
  • Reading date — The date recorded against every claim in this register, what it does and does not guarantee, and why language lists need it more than most claims.
  • Declaration — The difference between a claim a page makes to a reader and one it makes to a machine, and why this register counts the second kind.
  • Same-host rule — The rule this register uses to count site languages, the two alternatives, and what each of them would do to the comparison.
  • Unit mismatch — Comparing figures that count different things, the four units language figures come in here, and why a single column of numbers would be wrong.
  • Coverage table — The common comparison table of language counts, the four things it usually gets wrong, and which of those four are avoidable editorial choices.
  • Spot check — Opening a single declared address to see whether a declaration corresponds to a real translation, and why the result is never generalised.
  • Inclusion rule — The stated test for appearing in a register table, why it has to be written beside every table, and what it excludes here.
  • Dialect — A regional or social variety of a language, why counting dialects alongside languages changes a figure so much, and which vendors here keep them apart.
  • Accent — How somebody sounds rather than which words they use, why accent counts inflate figures further than dialect counts, and what an audience actually notices.
  • Variant — The standardised national form of a language, the four languages where it matters most in this register, and how rarely any vendor names one.
  • Mutual intelligibility — Whether speakers of two varieties understand each other, why that does not settle how many languages there are, and what it does to a count.
  • Writing system — The script a language is written in, the languages where more than one is in use, and why it is a rendering question as much as a language one.
  • Right-to-left text — Scripts that run right to left, what they change in an interface and in a picture, and how much of that any vendor here addresses.
  • Diacritic — Accents and marks above or below a letter, the languages where dropping one changes the meaning, and why generated pictures fail on them first.
  • Tonal language — Languages where pitch distinguishes words, why a synthetic voice can be intelligible and wrong, and what no vendor here publishes about it.
  • National standard — The codified form of a language used in schooling and broadcasting, why some languages have several, and how that maps onto a region subtag.
  • Voice inventory — The set of synthetic voices a product offers, why a language count is not an inventory figure, and what a series needs that neither states.
  • Speech synthesis — Producing speech from text, why a synthesis inventory is always smaller than a recognition one, and what that predicts about a vendor list.
  • Speech recognition — Turning recorded speech into text, why its language list is the longer of the two, and why the longer list is not what a buyer purchases.
  • Source language — The language of the material going into a tool, why it is a separate list from the output, and which questions it settles.
  • Target language — The language a finished output is delivered in, why it is the shorter list, and why a vendor quoting one number is usually quoting the other.
  • Dubbing — Replacing the dialogue of a finished video, what it can and cannot fix, and why a dubbing list is a different purchase from a generation list.
  • Lip sync — Adjusting mouth movement to match a new soundtrack, why it is a second pass rather than a property of a language list, and what makes it visible.
  • In-shot audio — Audio produced together with the picture rather than laid over it, what that removes from a pipeline, and what it does not promise.
  • Subtitle — Translated text laid over a finished picture, why it is not a voice claim at all, and the published limits that constrain what it can say.
  • Burned-in text — Text rendered permanently into the picture, why it is common in vertical short video, and what it costs a release in several languages.
  • Text inside the frame — Letters and characters a model draws inside the picture, why neither subtitles nor dubbing can reach them, and the one claim about it in this register.
  • Reading speed — How fast a viewer can read a subtitle, why the limit differs by language and audience, and why a vertical frame binds first.
  • Text expansion — The tendency of a translated sentence to be longer than the original, where the extra length accumulates, and which layers it breaks.
  • Transcreation — Adapting names, references and settings for a market rather than translating them, why it belongs at the script stage, and what it costs later.
  • Interface string — The individual pieces of text in a product interface, why translating them is the first language investment a company makes, and what it does not imply.
  • Working document — The outlines, breakdowns and descriptions a generated pipeline produces for people to act on, and why their language decides whether a tool is usable.
  • Market version — A release adapted for one market rather than translated, the layers it touches, and why a language count describes only one of them.
  • Localisation — The umbrella word for adapting a product or a release to a place, the layers it covers, and why any claim using it has to be split first.
  • Internationalisation — Building a product so that languages can be added later, why it is invisible from outside, and what its absence looks like.

A term is added when a register cell would otherwise have to explain itself, and never to describe a feature of one product.

The rest of the atlas: Columns, Entries, By language, Side by side, Markets, Tools, Output, Questions, Learn, Data. How a list earns a column is set out on the counting page.