TIL: UTF-8 ist eindeutig, Unicode nicht.
"ü" gibt es zweimal: U+00FC (NFC, C3 BC) und u + U+0308 (NFD, 75 CC 88). Gleiches Aussehen, aber verschieden.
Mein Importskript übernahm Buchtitel aus dem DNB-Katalog in NFD und hängte ein getipptes "- Hörbuch" in NFC an. Ein String, zwei Formen. Obsidian löst Links offenbar in NFC auf, und das Cover blieb leer.
Moral: Alles, was von außen kommt zuerst normalisieren.
unicodedata.normalize("NFC", s)
#Unicode #Python #Obsidian
#unicode
14 posts · Last used 11d
EBCDIC is incompatible with GDPR
https://shkspr.mobi/blog/2021/10/ebcdic-is-incompatible-with-gdpr/
Welcome to acronym city!
The Court of Appeal of Brussels has made an interesting ruling. A customer complained that their bank was spelling the customer's name incorrectly. The bank didn't have support for diacritical marks. Things like á, è, ô, ü, ç etc. Those accents are common in many languages. So it was a little surprising that the bank didn't support them.
The bank refused to spell their customer's name correctly, so the customer raised a GDPR complaint under Article 16.
The data subject shall have the right to obtain from the controller without undue delay the rectification of inaccurate personal data concerning him or her.
Cue much legal back and forth. The bank argued that they simply couldn't support diacritics due to their technology stack. Here's their argument (in Dutch - my translation follows)
Bank X also explained that the current customer data management application was launched in 1995 and is still running on a US manufactured mainframe system.
This system only supported EBCDIC ("extended binary-coded decimal interchange code"). This is an 8-bit standard for storing letters and punctuation marks, developed in 1963-1964 by IBM for their mainframes and AS/400 computers. The code comes from of the use of punch cards and only contains the following characters…
(Emphasis added.)
EBCDIC is an ancient (and much hated) "standard" which should have been fired into the sun a long time ago. It baffles me that it was still being used in 1995 - let alone today.
Look, I'm not a lawyer (sorry mum!) so I've no idea whether this sort of ruling has any impact outside of this specific case. But, a decade after the seminal Falsehoods Programmers Believe About Names essay - we shouldn't tolerate these sorts of flaws.
Unicode - encoded as UTF-8 - just works. Yes, I'm sure there are some edge-cases. But if you can't properly store human names in their native language, you're opening yourself up to a lawsuit.
Source
GDPRhub - 2019/AR/1006
DanceReactions
Marie ʕʘᴥʘʔ Julien
@mariejulienTrès intéressant !
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 40💬 2🔁 011:54 - Wed 20 October 2021
Grumpy Nat 🇨🇭🇧🇷🇲🇫
@Nat_KeelyHâte de mettre en justice tous les sites et autres compagnies qui ont décidé que le fait que j'ai un accent dans mon nom de famille soit source de bug (avec évidemment un message d'erreur qui n'a rien à voir. Histoire de bien pas comprendre pourquoi ça marche pas)
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 3💬 0🔁 010:43 - Wed 20 October 2021
Lays Y. M. Farra
@LYMFHSRLa France va sortir de l'UE juste pour que leur état-civil et autres administrations puissent continuer à ruiner la vie de quelqu'un parce qu'il a un tilde dans son nom
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 9💬 0🔁 013:38 - Wed 20 October 2021
KristoferA 🌏
@KristoferADoes this mean that Z̷̡̧̢̰͓̪͖̭͙̰̣̱̬̹̙̜̪̣̏̿̏̋͑́̒͑́̒̿̇̈̍̇̌͝͝a̵̡̧͍̘̮̤̙̹͙̦̙͙͖͓̥̟̦͔͒̇̊̊̔̓́͒́̌̈́̑͋̏̏̏̚͘͝͠͝l̶͉̯̱͇̭̭̉̉̈́̿͐̽̒̎̽͌̚͜ģ̸̧̛͙̩̹̰̤̱̖̘̻̪̻̮̫̟̙̲͍̰̻͕̗̫̿̆̃́͗̽̊̽̌̔̂͂̈͊̐̈́̈̈́̈̓̆͌̑́̕͜ǫ̶̢̹̥̮̟͍̔̑̔̽ can finally open a bank account?
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 16💬 0🔁 013:10 - Wed 20 October 2021
Bastien Nocera
@hadessukNext up, I’m suing La Poste for still using ISO-8859-1 when printing labels. Poor “Frédéric” I recently sent a game to…
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 7💬 0🔁 016:08 - Wed 20 October 2021
Michael Büker 🇺🇦
@emtiuEine Erschütterung der Macht, als würden Millionen Banken-ITler in panischer Angst aufschreien und dann verstummen.
Terence Eden is on Mastodon
@edentThis is interesting.
A bank claimed it couldn't use diacritics in a customer's name due to technical limitations.
Customer sued… and won!
Your name is personal data, and GDPR says it should be recorded accurately.❤️ 1,143💬 0🔁 49407:53 - Wed 20 October 2021❤️ 25💬 2🔁 006:13 - Thu 21 October 2021
#gdpr #name #unicode
Version 18 is published:
The #Unicode Blog: Announcing the Unicode® Standard, Version 18.0
https://blog.unicode.org/2026/09/announcing-unicode-standard-version-180.html?m=1
(Yeah, yeah, I know, I'm still at 16...)
I've run into a weird #Unicode Nomalization problem with #WordPress.
Thoughts and feedback welcome - https://core.trac.wordpress.org/ticket/66093
Unicode Variation Selector-15 and some of my tears
https://benjaminwil.info/weblog/variation-selector-15/
#Programming #Unicode #WebDev
Unicode strings that look the same aren't always the same. The @w3c@w3c.social's "Character Model for String Matching" document explains how case mapping, case folding, #emoji and other #Unicode details affect #interoperability and why specifications should define matching rules carefully #timetogiveinput #FPWD
▶️ https://www.w3.org/TR/charmod-norm/
The goal of this document is to provide a common reference for specification authors, software #developers, and content creators.
Feedback wlc: https://github.com/w3c/charmod-norm/issues
17 July is the one day of the year when the calendar #emoji is correct!
The word emoji combines the Japanese e (絵, picture) + moji (文字, character) and dates back to the late 80s. Its similarity to the English "emotion" is just a coincidence.
Nowadays, emoji are standardised by the Unicode Consortium, a Californian non-profit that has its roots in facilitating a universal digital representation of the world's writing systems. #Unicode doesn't specify exactly how emoji should look – so they're different on Android and Apple devices, for example – but it lists their names.
The first major character set used in computing was #ASCII, the American Standard Code for Information Interchange, published in 1963. Using 7 bits, it could represent 128 characters — more than enough for a to z, A to Z, 0 to 9 and common punctuation marks, but certainly indicative of a North American worldview.
Later "extended" versions of ASCII doubled this to 256 possibilities, allowing for dozens of accented letters, characters like æ and ß, and symbols like © and ¾. And double-byte character sets like Shift JIS meant that Chinese, Korean and Japanese could be represented in full. But multiple standards existed, and it became just a bit too common to see ����.
The answer was Unicode: a single, multi-byte character encoding standard for everyone. Plus emoji!
#language #i18n
Replying to
@0xabad1dea@infosec.exchange
Yesh, that's definitely what you do if you find a random snake in a hotel.
Stuff it in your bra and try to sneak it through airport security/customs.
Hey, #Unicode geeks, where's the facepalm emoji?
Here's How to Officially Suggest New Emojis
https://lifehacker.com/tech/how-new-emojis-are-created?utm_medium=RSS
#Tech #Emoji #Unicode
a nice unexpected benefit of godbolt.org is that it's making me aware of more system languages.
today, i learned about the C3 language via the godbolt language pull-down menu.
for prototyping purposes (not for production, which requires memory safety), C3 seemed to check a lot of boxes.
but then i got to the part where it milkshake-ducks itself:
https://c3-lang.org/faq/rejected-ideas/#unicode-identifiers
they might as well have asked, "why would anyone think in their native language, even in their personal experimental code?"
The footgun of right-to-left decorative characters
https://blog.alexbeals.com/posts/the-footgun-of-right-to-left-decorative-characters
#HackerNews #Tech #Unicode
Unicode's Transliteration Rules Are Turing-Complete
https://seriot.ch/computation/uts35/
#ComputerScience #Unicode #Programming
Replying to
@Profpatsch@mastodon.xyz You "added #Unicode username support to #flohmarkt" last month — wonderful to hear! I will look for your demo at #FediForum. I will propose a session on "#Globally-inclusive Fediverse handles". If you were present, it would improve the session. #UniversalAcceptance #Multilingual
Replying to
You've seen all posts