FaceDeer
FaceDeer@fedia.io
<p>Basically <a href="https://imgur.com/9Wg5TTx.png">a deer with a human face</a>. Despite probably being some sort of magical nature spirit, his interests are primarily in technology and politics and science fiction.</p> <p>Spent many years <a href="https://www.reddit.com/user/FaceDeer/">on Reddit</a> before joining the Threadiverse as well.</p>
Posts
-
Post #4143633
It's a 5km diameter target area. It's not going to be shining on random patches of wilderness.
-
Post #638327
My freakout would be "oh, this is so cool!"
-
Post #444716
The companies that continue with human staff where others are replacing theirs, for example. Outsourcers providing those staff.
-
Post #441013
No, we fight to ensure that the rules are followed. In this case they are, the judge has discretion here. Would you rather there were "mandatory minimum" laws when it came to this as well?
-
Post #440842
Whereas I prefer an organized rules-based justice system over anarchy and vigilantism. Because who knows when you or I might end up being in the "disliked" category?
-
Post #439388
But don't you see? We don't like these particular people, so they should suffer the maximum possible penalties under every circumstance. If we liked them then punishing them for wearing glasses would of course be a travesty.
-
Post #438693
Depends entirely on the circumstances of where it does end up being built. I'm not sure what "gotcha" you think you're making here? That Reddit comment is just me pointing out that when a business uses electricity they pay for it.
-
Post #231305
It's weird how AI has turned so much of the internet from its generally anti-copyright stance. I've seen threads in piracy and datahoarding communities that were riddled with "won't someone please think of the copyright!" Posts raging about how awful AI was. I maintain the same view I always have. Copyright is indeed broken, because of how overly restrictive and expansive it has become. Most people long ago lost sight of what it's actually for.
-
Post #231298
If they do it's not by the actual training of AI.
-
Post #230229
Simple.wikipedia isn't a summary of regular Wikpedia, it's a whole separate thing. It's intended to convey the same data, just in a simpler way.
-
Post #230228
The problem being discussed here is not the availability of Wikipedia's data. It's about the ongoing maintenance and development of that data going forward, in the future. Having a static copy of Wikipedia gathering dust on various peoples' hard drives isn't going to help that.
-
Post #230225
Wikipedia's traditional self-sustaining model works like this: Volunteers (editors) write and improve articles for free, motivated by idealism and the desire to share knowledge. This high-quality content attracts a massive number of readers from search engines and direct visits. Among those millions of readers, a small percentage are inspired to become new volunteers/editors, replenishing the workforce. This cycle is "virtuous" because each part fuels the next: Great content leads...
-
Post #192933
I like that one, very thorough.
-
Post #192878
LLMs were trained on our social media feeds, after all...
-
Post #148800
Which, as I said, seems strange. Why don't those businesses just download the torrents?
-
Post #142586
Youtube isn't the way you think it should be, though.
-
Post #142148
"Why did they take that feature away? I was busy abusing it!"
-
Post #132423
Those would be the Epstein Files. This is still nice to have out there, though. Governments should fear their people.
-
Post #129831
I have a sneaking suspicion that the vast majority of the people raging about AIs scraping their data are not raging about it being done inefficiently.
-
Post #129828
You're thinking of "model decay", I take it? That's not really a thing in practice.
-
Post #129597
Raw materials to inform the LLMs constructing the synthetic data, most likely. If you want it to be up to date on the news, you need to give it that news. The point is not that the scraping doesn't happen, it's that the data is already being highly processed and filtered before it gets to the LLM training step. There's a ton of "poison" in that data naturally already. Early LLMs like GPT-3 just swallowed the poison and muddled on, but researchers have learned how much bet...
-
Post #129555
I have no idea what "established means" would be. In the particular case of the Fediverse it seems impossible, you can just set up your own instance specifically intended for harvesting comments and use that. The Fediverse is designed specifically to publish its data for others to use in an open manner.
-
Post #129500
Are you proposing flooding the Fediverse with fake bot comments in order to prevent the Fediverse from being flooded with fake bot comments? Or are you thinking more along the lines of that guy who keeps using "Þ" in place of "th"? Making the Fediverse too annoying to use for bot and human alike would be a fairly phyrric victory, I would think.
-
Post #129104
A basic Google search for "synthetic data llm training" will give you lots of hits describing how the process goes these days. Take this as "defeatist" if you wish, as I said it doesn't really matter. In the early days of LLMs when ChatGPT first came out the strategy for training these things was to just dump as much raw data onto them as possible and hope quantity allowed the LLM to figure something out from it, but since then it's been learned that quality is bett...
-
Post #128639
I think it's worthwhile to show people that views outside of their like-minded bubble exist. One of the nice things about the Fediverse over Reddit is that the upvote and downvote tallies are both shown, so we can see that opinions are not a monolith. Also, engaging in Internet debate is never to convince the person you're actually talking to. That almost never happens. The point of debate is to present convincing arguments for the less-committed casual readers who are lurking rather t...
-
Post #85060
Ah, good, that makes this less of a dilemma then.