Elektrine lite

← Feed

@varx@infosec.exchange

2026-07-28 02:08 UTC

@gaditb@icosahedron.website Yeah, it was inevitable that we would eventually come to an encoding that would allow representing most/all characters from world languages, unambiguously. The previous situation is untenable. However, I wonder if you would call the following "a unicode": A container format for text where there is a prefix indicating the character set (and maybe some delimiters, runlength indicators, whatever). A multi-lingual plaintext document would have different parts in different character sets.

Replies (1)

  • @gaditb@icosahedron.website 2026-07-28 05:52

    @varx@infosec.exchange I would distinguish that. The reason I would distinguish that from a Unicode is that it: a) leaves the character set as an open category (you could argue that the fundamental extensibility of Unicode and/or the private-use areas enable that, but that feels different to me, more all-encompassing. I dunno, maybe I'm wrong.) b) doesn't inherently priviledge the non-diacritic-having Latin alphabet in the encoding (or at least, only minimally, using the first available prefix specifically) (You could again argue that the same thing applies to Unicode, that it is simply being at the beginning of Plane 0 that is the extent of that priviledging, but first of all that often means a 2+-times difference in size required to encode a single character and second of all often non-Latin scripts are often split into muliple discontinuous regions.)

    Open ##4645159