The vocabulary a language list is written in
These pages define the words a language claim is made of. None of them contains a figure about a named product, because a definition that depends on one vendor is not a definition. As of 2026-09-12.
Almost every argument about language coverage is a vocabulary problem. One word covers a translated menu, a generated script and a synthetic voice, and a sentence using it can be true of one and false of the others without changing a character.
So each term is given its own page with the same three parts: what it means, what it pins down, and what readers take it to promise that it does not. Where a term has a narrow technical sense and a loose marketing sense, both are written down and the gap between them is the point.
These pages are the layer the register leans on. A cell can stay short because the distinction it depends on is defined once here instead of being restated in a parenthesis on every row.
1The vocabulary
- Language tag — The complete code a page or a voice is labelled with, which parts of it are required, and how much a reader can infer from the parts that are missing.
- Language subtag — The two or three letters naming a language, why some languages have two codes, and what a comparison of bare codes can and cannot conclude.
- Region subtag — The country or region appended to a language code, what it pins down for a reader, and why most of the vendors in this register leave it off.
- Script subtag — The part of a code that names a writing system, the few languages where it is necessary, and why it is not interchangeable with a region.
- hreflang attribute — The markup that tells a search engine which address serves which language, why this register counts it, and the four questions it cannot answer.
- x-default — The alternate entry that names no language, what it is for, and why counting it would measure configuration quality instead of language coverage.
- lang attribute — The attribute in which a document states its own language, how it differs from an alternate declaration, and what it is worth as evidence.
- Locale — The bundle of language, region and formatting conventions a product switches between, and why a locale count is not a language count.
- Language cluster — A group of addresses declaring each other as language versions of one page, why vendors maintain them, and what a broken cluster looks like.
- Language switcher — The menu a visitor uses to change a site language, why it is not what this register counts, and what it can still tell a reader.
- Machine-swapped page — A page put through automatic translation and published unreviewed, how to recognise one, and why a declared alternate cannot rule it out.
- Language subdomain — Serving each language from its own subdomain, why the counting rule discards those alternates, and what the discarded figure still tells a reader.
- Country-code domain — A separate national domain for a market, why an alternate pointing at one is not counted here, and what cannot be read from the markup.
- Named list — A published set of language names, why it is worth more than any count, and the four forms it takes among the vendors read here.
- Bare count — A language figure published without the set it counts, why the size of the number is irrelevant, and why the register still records it.
- Plus figure — The habit of writing a language count with a plus after it, what that does to the claim, and why it removes the pressure to keep a page current.
- Printed total — A list carrying its own count, why it is the most useful shape of disclosure in this register, and how rarely any vendor bothers with it.
- Per-job cap — A limit on how many languages a single job can produce, why it affects a schedule more than an inventory figure does, and how rarely it is published.
- Empty cell — What a blank in this register asserts, what it deliberately does not, and why the distinction changes how a whole column should be read.
- Reading date — The date recorded against every claim in this register, what it does and does not guarantee, and why language lists need it more than most claims.
- Declaration — The difference between a claim a page makes to a reader and one it makes to a machine, and why this register counts the second kind.
- Same-host rule — The rule this register uses to count site languages, the two alternatives, and what each of them would do to the comparison.
- Unit mismatch — Comparing figures that count different things, the four units language figures come in here, and why a single column of numbers would be wrong.
- Coverage table — The common comparison table of language counts, the four things it usually gets wrong, and which of those four are avoidable editorial choices.
- Spot check — Opening a single declared address to see whether a declaration corresponds to a real translation, and why the result is never generalised.
- Inclusion rule — The stated test for appearing in a register table, why it has to be written beside every table, and what it excludes here.
- Dialect — A regional or social variety of a language, why counting dialects alongside languages changes a figure so much, and which vendors here keep them apart.
- Accent — How somebody sounds rather than which words they use, why accent counts inflate figures further than dialect counts, and what an audience actually notices.
- Variant — The standardised national form of a language, the four languages where it matters most in this register, and how rarely any vendor names one.
- Mutual intelligibility — Whether speakers of two varieties understand each other, why that does not settle how many languages there are, and what it does to a count.
- Writing system — The script a language is written in, the languages where more than one is in use, and why it is a rendering question as much as a language one.
- Right-to-left text — Scripts that run right to left, what they change in an interface and in a picture, and how much of that any vendor here addresses.
- Diacritic — Accents and marks above or below a letter, the languages where dropping one changes the meaning, and why generated pictures fail on them first.
- Tonal language — Languages where pitch distinguishes words, why a synthetic voice can be intelligible and wrong, and what no vendor here publishes about it.
- National standard — The codified form of a language used in schooling and broadcasting, why some languages have several, and how that maps onto a region subtag.
- Voice inventory — The set of synthetic voices a product offers, why a language count is not an inventory figure, and what a series needs that neither states.
- Speech synthesis — Producing speech from text, why a synthesis inventory is always smaller than a recognition one, and what that predicts about a vendor list.
- Speech recognition — Turning recorded speech into text, why its language list is the longer of the two, and why the longer list is not what a buyer purchases.
- Source language — The language of the material going into a tool, why it is a separate list from the output, and which questions it settles.
- Target language — The language a finished output is delivered in, why it is the shorter list, and why a vendor quoting one number is usually quoting the other.
- Dubbing — Replacing the dialogue of a finished video, what it can and cannot fix, and why a dubbing list is a different purchase from a generation list.
- Lip sync — Adjusting mouth movement to match a new soundtrack, why it is a second pass rather than a property of a language list, and what makes it visible.
- In-shot audio — Audio produced together with the picture rather than laid over it, what that removes from a pipeline, and what it does not promise.
- Subtitle — Translated text laid over a finished picture, why it is not a voice claim at all, and the published limits that constrain what it can say.
- Burned-in text — Text rendered permanently into the picture, why it is common in vertical short video, and what it costs a release in several languages.
- Text inside the frame — Letters and characters a model draws inside the picture, why neither subtitles nor dubbing can reach them, and the one claim about it in this register.
- Reading speed — How fast a viewer can read a subtitle, why the limit differs by language and audience, and why a vertical frame binds first.
- Text expansion — The tendency of a translated sentence to be longer than the original, where the extra length accumulates, and which layers it breaks.
- Transcreation — Adapting names, references and settings for a market rather than translating them, why it belongs at the script stage, and what it costs later.
- Interface string — The individual pieces of text in a product interface, why translating them is the first language investment a company makes, and what it does not imply.
- Working document — The outlines, breakdowns and descriptions a generated pipeline produces for people to act on, and why their language decides whether a tool is usable.
- Market version — A release adapted for one market rather than translated, the layers it touches, and why a language count describes only one of them.
- Localisation — The umbrella word for adapting a product or a release to a place, the layers it covers, and why any claim using it has to be split first.
- Internationalisation — Building a product so that languages can be added later, why it is invisible from outside, and what its absence looks like.
A term is added when a register cell would otherwise have to explain itself, and never to describe a feature of one product.
The rest of the atlas: Columns, Entries, By language, Side by side, Markets, Tools, Output, Questions, Learn, Data. How a list earns a column is set out on the counting page.