Elektrine lite

← Feed

@Azuaron@cyberpunk.lol

2026-09-18 00:34 UTC

@mttaggart@infosec.exchange There are a bunch of separate things at play, some of which are pretty cut and dry, some are less so. tldr; I think Microsoft/OpenAI are facing several violations: Unauthorized access of the NYT servers (computer fraud)Illegal acquisition of NYT articles (copyright)Potentially infringing LLM output (copyright)Violation of the Hot News Doctrine (quasi-property rights) The first part is "how did they acquire the content." Depending on the method of their "hack", they might be guilty of not a copyright claim, but the Computer Fraud and Abuse Act for unauthorized access of a computer system. Of course, maybe they get hit for both some kind of hacking and also the copyright. Anthropic got in trouble for illegal acquisition, but their acquisition was via torrenting, which is a legally well-defined copyright violation. It's unclear to me how "hacked a computer system" would make the illegal copyright acquisition go away, so it seems like Microsoft/OpenAI should get hit for the copyright issue here just like Anthropic did. And, of course, the judge in the Facebook case went full insane and said torrenting was legal for Facebook, so anything's possible, I guess. After the acquisition, we've got the training. "Training" is generally considered to be Fair Use, in and of itself, since "training" really just means "collecting mathematical statistics on word sequences", and we have decades of legal precedent saying that you can collect mathematical facts regarding copyrighted material and it's Fair Use. Finally, there's the question of distribution. Generally speaking, the question of LLM distribution usually doesn't reach the Fair Use question, because it doesn't pass the "substantial similarity" test: something is only a copyright violation if it is "substantially similar" to the copyrighted work, and LLM output is usually not similar enough to any individual copyrighted work for LLM output to be a copyright violation. However, that's generally, so let's talk about a case out of Germany. There was a music AI company that musicians claimed was outputting substantially similar work, and the AI company claimed that it was because the prompters were "steering" the AI to generate copyrighted works, so the prompters should be liable. The judge's ruling came down and it did ultimate hinge on that steering question: if the prompters could be said to be steering the AI to reproduce copyrighted work, then the prompters would be responsible for the violation, but if the prompters were asking for something generic and the AI was providing substantially similar music as output, then the AI company would be liable (in this case, the judge ruled that the AI company was liable). It should be further noted: the judge in the Anthropic case specifically said if Claude was producing substantially similar output, then Anthropic would be liable for that. So it seems pretty likely that if the LLM is reproducing copyrighted work, and the prompters aren't "steering" it, the AI company is liable for copyright infringement. Apart from the steering, though, why was the AI outputting substantially similar music? Well, in the case of word-based LLMs, they're drawing on trillions and trillions of inputted documents. There's so much less music. So, if you prompt for a less-than-popular genre and a couple of "with this kind of beat" or whatever, the AI doesn't have a lot of music to draw on for its response, so often ends up outputting music that's substantially similar to copyrighted music. In the NYT case, I don't think they're going to get away with the "prompter is steering it"; people are going to put in prompts like "what's going on with ____" or "tell me the latest news", and they're going to get information straight out of a small number of publications, of which NYT is a significant part. Since all the information is coming out of a small number of publications, the AI companies are going to have a similar problem as the music AI: there's only so much mixing the LLM can do with just a few sources, and the output's likely to be substantially similar to at least one of their sources. But even if the LLM output is not substantially similar, there is still the Hot News Doctrine. Essentially, during WWI the AP was gathering news from the war and publishing it on the east coast, and the International News Service would take those articles, telegraph them to the west coast, and beat the AP to publication over there. The courts ruled (and have affirmed several times, including recently) that even though facts were not copyrightable, the AP had "quasi-property rights" over the facts that they had expended a lot of time and energy finding and determining, and the INS' free-riding on the AP's hard work constituted unfair competition. There is a five element test for determining a violation of the Hot News Doctrine: 1) A plaintiff generates or gathers information at a cost; 2) The information is time-sensitive; 3) A defendant's use of the information constitutes free riding on the plaintiff's efforts; 4) The defendant is in direct competition with a product or service offered by the plaintiffs; 5) The ability of other parties to free-ride on the efforts of the plaintiff or others would so reduce the incentive to produce the product or service that its existence or quality would be substantially threatened. It seems incredibly likely to me that the AI companies are violating the Hot News Doctrine. They're literally admitting point 5, which seems to be the point of most contention in previous cases. Indeed, last year, in this very lawsuit, the judge throw out a claim by NYT regarding the Hot News Doctrine because OpenAI was only training on the articles, not, in a timely way, distributing the reporting. I don't think that's still the case; if ChatGPT is outputting timely news that it didn't get a license for, the judge explicitly said that while training was not Hot News freeriding, scraping and then turning around and outputting LLM news of the scraped information could be. I am, also, not a lawyer, I just have a hobbyist's interest in this kind of thing, so grain of salt and all that.

Replies (0)

No replies.