Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Why? What is the goal of a scraper, and how does disabling the source of the data benefit them?


> Why? What is the goal of a scraper, and how does disabling the source of the data benefit them?

The next scraper doesn’t get the data. People don’t realize we’re not compute limited for ai, we’re data limited. What we’re watching is the “data war”.


at this point we're _good data_ limited, which has little to do with scraping.


Why kind of data that isn’t public would be so valuable for AI training?

Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc.

I can understand targeting certain sites like Reddit, etc. but not random websites


It's to rip off copyrighted content and profit from it instead of the original authors. It's like every other low rent and highly automated scam that finds it's way onto the internet.

If you look closely even Google does this. This is probably why many popular sites started getting down ranked in the last 2 years. Now they're below the fold and Google can present their content as their own through the AI box.


Please remember that Google only needs to be marginally better than the competition. And, of course, their primary biz is ads, not serving great results; that is a distant second priority.


Their biz is ads, but since search is winner takes all they need only be marginally better than the competition... twenty years ago.


> Their biz is ads,

Yea, but, the FTC doesn't want it to be.


Discord I guess would be quite valuable, even the de facto public servers.


Honestly it's hard to tell how much more value the LLM people are going to get out of another copy of the internet.

It feels a lot like they're stuck for improvements but management doesn't want to hear it.


It's a bit strange to talk about stuck when the most recent breakthrough is less than a year old.


I’m not sure what you mean by breakthrough, but if you’re talking about Deepseek, it’s more of an incremental improvement than a breakthrough.


Scraping social media is good data, even without ML. The fact that something is "happening" to people in a social space inherently has importance to people. The specter of law is more threatening to whether companies can get their hands on good data.


Now the only way to obtain that information is through them


I guess one could make a point that competition will no longer have the access to the scraped data.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: