2026-09-18 11:17 UTC
Replies (1)
-
@masek@infosec.exchange 2026-09-18 11:17
Training also requires content. The early norm was often: if material was accessible online, it was available for ingestion. Whether that is lawful remains contested across jurisdictions. Whether it was fair to creators is a separate question. Now many actors repeatedly crawl the same web. This imposes costs on the people and institutions that produced the value. Wikimedia reported that bots generated at least 65% of its most resource-intensive traffic. Other infrastructure providers document AI crawlers ignoring restrictions or obscuring their identity. Not every automated request is for training. Some support search or user queries. But the larger pattern remains: private acquisition pipelines duplicate load while shifting bandwidth, hosting, moderation, and defense costs onto the open web. The web is treated as a buffet, except the bill is sent to the people who cooked. That is a poor foundation for a public-benefit narrative. 12/34