Post #2542018
2026-05-14 11:15 UTC
While preparing for the next session of my LLM class on training data, I came across this brilliantly illustrated article from the @washingtonpost analysing the content of Google’s C4 data set, a filtered version of the Common Crawl used as training data for many LLMs: https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/ (free access). Added to the seminar's required reading list!
#dataviz #LLM #GenAI
Replies (1)
-
@Robotistry@fediscience.org 2026-05-14 12:27
@ElenLeFoll@fediscience.org @washingtonpost@mstdn.social Charming how it ends by noting that "Social networks like Facebook and Twitter — the heart of the modern web — prohibit scraping, which means most data sets used to train AI cannot access them." Would these companies (who are described in the article as stealing copyrighted content and whose models have been reported elsewhere as identifying enormous numbers of security vulnerabilities) really be deterred by a prohibition?