Language tag: the whole string, and how much of it is optional
A language tag is the whole code: a required language part, and optional script, region and further parts after it. Almost every tag in this register is the language part alone, which is legal and which answers one question rather than three. As of 2026-09-22.
| Part | Required | What it adds |
|---|---|---|
| Language | Yes | Which language the content is in |
| Script | No | Which writing system, where more than one is used |
| Region | No | Whose national norms the wording follows |
| Variant and extensions | No | Finer distinctions, absent from everything read here |
Inclusion rule. The parts of a tag that appear in the material this register reads. Extensions and private-use parts are listed for completeness only. Order. In the order the parts appear in a tag.
1One required part and three optional ones
The shortest legal tag is two or three letters naming a language. Everything after that narrows it: a script, then a region, then finer variants. A tag can stop at any point, and most of the tags in this register stop immediately.
Stopping early is not sloppiness. A bare tag is the correct choice when a page is meant for every reader of a language, and it becomes a problem only when the thing being labelled differs between markets, which is true of speech far more than of written pages.
2Reading a tag backwards is a mistake
A longer tag is more specific, not better. Naming a country rules markets out as well as in, and a vendor that tagged every page to one country would be making a narrower offer than one that left them open.
So the register records the form beside the count instead of ranking forms. Two cells reading the same number can hold tags of different lengths, and which of those a reader prefers depends on whether they want to be included or to be sure.
3Why the string is transcribed rather than normalised
Tags could be reduced to their language part before being stored, which would make counting simpler and would delete the only variant evidence these pages carry. A cell saying one thing and a cell saying another are different facts.
Counting is done on the language part, so a vendor declaring one regional variant counts once for that language. The full string is printed, so a reader who cares about the variant can see it without trusting the count.
A definition, not a measurement: no vendor figure appears here that is not also in the register. The parts are defined separately under language, script and region subtag.
Nearby terms: Internationalisation, Language subtag. The whole vocabulary is at terms; nothing on this page is a claim about a named product.