2026-07-28 00:20 UTC
Okay if I'm understanding people's responses to the above correctly, I expect this to get an overwhelming (uniform? maybe not?) "Yes":
Was a universal encoding that would see a nontrivial sequence of bytes as translating unambiguously to a particular sequence of characters, selected out of a set intended to contain all* in-use characters across humanity,
inevitable?
(*for practical purposes. E.g. modulo CJK unification, okay skipping until-sufficiently-popular e.g. sitelen pona and other small-use symbols. Also okay skipping (or not) in the hypothetical cases emoji, etc.)
This has an open-ended follow-up question :...
Replies (1)
-
@varx@infosec.exchange 2026-07-28 02:08
@gaditb@icosahedron.website Yeah, it was inevitable that we would eventually come to an encoding that would allow representing most/all characters from world languages, unambiguously. The previous situation is untenable. However, I wonder if you would call the following "a unicode": A container format for text where there is a prefix indicating the character set (and maybe some delimiters, runlength indicators, whatever). A multi-lingual plaintext document would have different parts in different character sets.