Elektrine lite

← Feed

@axbom@axbom.me

Post #1868939

2023-10-22 08:19 UTC

Here's what happens when machine learning needs vast amounts of data to build statistical models for responses. Historical, debunked data makes it into the models and is preferred by the model output. There is much more outdated, harmful information published than there is updated, correct information. Hence statistically more viable. "In some cases, they appeared to reinforce long-held false beliefs about biological differences between Black and white people that experts have spent years trying to eradicate from medical institutions." In this regard the tools don't take us to the future, but to the past. No, you should never use language models for health advice. But there are many people arguing for exactly this to happen. I also believe these types of harmful biases make it into more machine learning applications than language models specifically. In libraries across the world using the Dewey Decimal System (138 countries), LGBTI (lesbian, gay, bisexual, transgender and intersex) topics have throughout the 20th century variously been assigned to categories such as Abnormal Psychology, Perversion, Derangement, as a Social Problem and even as Medical Disorders. Of course many of these historical biases are part of the source material used to make today's "intelligent" machines - bringing with them the risk of eradicating decades of progress. It's important to understand how large language models work if you are going to use them. The way they have been released into the world means there are many people (including powerful decision-makers) with faulty expectations and a poor understanding of what they are using. https://www.nature.com/articles/s41746-023-00939-z #DigitalEthics #AIEthics

Replies (15)

  • @kegill@mastodon.social 2023-10-22 09:19

    @axbom Thank you. I will share this with my undergrad engineering students. I knew this (more bad or outdated than good) but had not made that next step.

    Open ##1868999

  • @riley@toot.cat 2023-10-22 09:23

    @axbom Somebody needs to invent a way to methylate obsolete and wrong ideas inside an LLM.

    Open ##1869000

  • @AT1ST@mstdn.ca 2023-10-22 10:52

    @axbom I'm reminded of an older lecture where someone explaining how they were testing their "Identify if this is a wolf or a dog" neural network...and found that it misidentified dogs as wolves if there was snow in the picture, and vice versa if there isn't snow. Historically, it's probably more likely that wolves pictures were taken in the snow, so that bias may make sense in hindsight...except that's not the only place that we want to identify them.

    Open ##1869002

  • @axbom be part of @theeyeofmasonic

    Open ##1869004

  • @axbom #medlibs scroll up for an interesting paper about biased, incorrect, pejorative, outdated training data in LLMs

    Open ##1869015

  • @CharlesMicah@tech.lgbt 2023-10-22 12:42

    @axbom SMBC had thoughts about this https://www.smbc-comics.com/comic/copyright

    Open ##1869024

  • @debrecenisrac46@masto.ai 2023-10-22 13:26

    @axbom when it comes to chatgpt it's scary how easily it falls for top search results in Google. https://divany.hu/vilagom/kossuth-lajos-szolidaritas/ I asked chatgpt about Kossuth and it gave a glowing speech based on mostly translated sources. When I confronted it that both in England and in the US he was critized for being silent on the treatment of the Irish or the slaves in America, it course corrected to reflect and tell me something I already knew.

    Open ##1869031

  • @axbom I've seen that kind of thing before. I worked quite a bit with a mathematical estimation technique called a Kalman filter. It would take in measurements of interest and modeling functions and put out modeling coefficients. Some managers thought we should just pour everything into the filter and like magic, answers would appear. I tried to explain the "garbage in garbage out" philosophy, that putting in measurements known to be bad wasn't useful. I was partially successful.

    Open ##1869037

  • @flux@derg.social 2023-10-22 13:49

    @axbom@axbom.me louder for the people in the back!! Great post

    Open ##1869042

  • @enigmatico@mk.absturztau.be 2023-10-22 14:04

    @axbom@axbom.me When did we start listening to machines generating text using an stochastic model predicting human speech with several hard coded biases, and thought it was a good idea? These machines won't say the truth. They will try to guess what a human would think it's the truth, and generate a text that looks like the truth.

    Open ##1869057

  • @scotty86@mastodon.social 2023-10-22 14:56

    @axbom The term '#AI' often misleads, suggesting a level of smartness that #ChatGPT doesn't possess. I view ChatGPT more as a language processor/advanced query system: it deciphers human language, navigates through data, and responds in human language. It's less about artificial 'intelligence' and more about advanced language and data processing.

    Open ##1869058

  • @fabiobrady@mastodonapp.uk 2023-10-22 15:08

    @axbom what about an LLM trained on a pile of up to date more recent research?

    Open ##1869059

  • @otto42@fosstodon.org 2023-10-22 18:31

    @axbom Of course, because they're trying to teach it to learn language, not to speak actual truth. That's why these models will lie to you, because they don't know any better, because they don't understand fiction from fact, and because they were never built to discern that.

    Open ##1869060

  • @CaseyWrites@toot.community 2023-10-22 22:13

    @axbom I’m both excited about AI and concerned about #AIethics. Every time #ChatGPT generates something awesome, I cheer. Then I hear Jeff Goldblum in my head. John Hammond : I don't think you're giving us our due credit. Our scientists have done things which nobody's ever done before... Dr. Ian Malcolm : Yeah, yeah, but your scientists were so preoccupied with whether or not they could that they didn't stop to think if they should.

    Open ##1869062

  • @woo@fosstodon.org 2023-10-23 00:18

    @axbom "90% of the Internet is wrong". "90% of everything is wrong". Sorry, I don't remember who said it.

    Open ##1869063