All of their business decisions for the past few years have been centered around monetizing Reddit content as training data for machine learning models. Most likely this recent decision is an attempt to block data scraping.
Even within the lemmy user community people still occasionally refer to reddit because there isn’t a better source for some things. Outside of lemmy, reddit is a household name alongside facebook, snapchat &etc; people who don’t use it still know what it is. I think your conclusion is more aspirational than currently real.
Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.
All of their business decisions for the past few years have been centered around monetizing Reddit content as training data for machine learning models. Most likely this recent decision is an attempt to block data scraping.
It’s gonna block sign ups. Why the fuck do I sign up without seeing shit?
Not like reddit is relevant these days.
Even within the lemmy user community people still occasionally refer to reddit because there isn’t a better source for some things. Outside of lemmy, reddit is a household name alongside facebook, snapchat &etc; people who don’t use it still know what it is. I think your conclusion is more aspirational than currently real.
The opacity is going to make the issue very real. A site like that goes stagnant without new blood.
Why anyone would train an AI on Reddit is beyond me, the only worse platform I can think of is 4chan.
Because it has years of helpful troubleshooting guides, solutions to problems, and conversations. also there is an llm trained off 4chan, gpt-4chan.
Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.
Everything posted in the last few years should be considered poison anyway, so I don’t know why anyone would pay for it to use as training data.