Post #1082576
2025-12-16 22:26 UTC
Replies (2)
-
@csolisr@hub.azkware.net 2025-12-17 15:42
Problem is, among the more famous models, even the most open ones are still trained partially on copyrighted content, such as the "Common Crawl" dataset. The problem with sourcing and attribution has not been sufficiently tackled by most LLM producers - in terms of sourcing, the most promising attempt I've found is the "Common Pile", a dataset comprised entirely of public domain and copyleft data sources, but nearly nobody is using it as a base. huggingface.co/papers/2506.052…
-
@ppw@social.redflag.ps 2025-12-18 01:50
@msokiovt@uwu.social @Waterfox@mastodon.social there is nothing wrong with having AI features in your browser, but it should be up to the user to decide whether and how this happens, rather than Firefox simply imposing it on you and then not even giving you the option to disable it.