

Admittedly, no, but this is slowly improving. Stepfun published their SFT dataset, smaller labs are publishing their datasets for task specific tunes. I believe there was another Chinese lab that published bulk pretraining data, but I can’t find it in my history at the moment.
And, notably, these comparatively tiny labs generally aren’t scraping the internet so abusively like OpenAI/Meta. They don’t need as much. Going by statements in their papers, they tend to use existing archives of web data, buy commercial data, or (more recently) generate a lot synthetic data.

Reddit seems to be highly algorithmic now, so I think it has more of the “doomscrolling effect” that does accrue a lot of engagement.