Post #1082577
2025-12-17 15:42 UTC
Problem is, among the more famous models, even the most open ones are still trained partially on copyrighted content, such as the "Common Crawl" dataset. The problem with sourcing and attribution has not been sufficiently tackled by most LLM producers - in terms of sourcing, the most promising attempt I've found is the "Common Pile", a dataset comprised entirely of public domain and copyleft data sources, but nearly nobody is using it as a base. huggingface.co/papers/2506.052…
Replies (0)
No replies.