Elektrine lite

← Feed

@masek@infosec.exchange

2026-09-18 11:17 UTC

Training also requires content. The early norm was often: if material was accessible online, it was available for ingestion. Whether that is lawful remains contested across jurisdictions. Whether it was fair to creators is a separate question. Now many actors repeatedly crawl the same web. This imposes costs on the people and institutions that produced the value. Wikimedia reported that bots generated at least 65% of its most resource-intensive traffic. Other infrastructure providers document AI crawlers ignoring restrictions or obscuring their identity. Not every automated request is for training. Some support search or user queries. But the larger pattern remains: private acquisition pipelines duplicate load while shifting bandwidth, hosting, moderation, and defense costs onto the open web. The web is treated as a buffet, except the bill is sent to the people who cooked. That is a poor foundation for a public-benefit narrative. 12/34

Replies (1)

  • @masek@infosec.exchange 2026-09-18 11:17

    My conclusion is blunt: the unchecked race of ever larger, privately controlled training bets has to stop. I do not mean research, every new model, or the use of LLMs. I mean that this cannot continue as normal. OpenAI and Anthropic appear to recognize at least part of the problem. Some industry leaders have called for pacing the frontier; OpenAI says it would slow or stop when risks cannot be sufficiently safeguarded. But they are riding a tiger, and tigers are unhelpful when asked for safe dismounting instructions. Slowing threatens valuation, funding, talent, market position, and the story justifying infrastructure commitments. Continuing deepens the dependencies. I see no credible off-ramp. If the companies cannot find one, politics and society must build it from outside. The brake cannot depend on the riders alone. We need public rules, coordination, transparency, and a way to absorb the economic consequences. 13/34

    Open ##4748050