The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Frankly, I’m surprised they still offer RSS feeds. They’ve been slowly but surely killing off all ways of accessing their content for years. One day they’ll disable them, but for now they still work.
so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.
I blocked the fuckers, no qualms at all.
Even if they were using the API they were not being nice about their shit.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
Its amazing to me how consistently the shitty behaviours of these billionaire techbro oligarchs impact disabled or marginalised people… even when the point isnt to directly shit on them. Its fucking vile.
I honestly think many (too many, but certainly not all! I am one) programmers are some of the immoral, ethically spurious people around in the 21st century.
That is why they want AI to replace programmers. AI morals are programmed, so they can be designed to do shitty things that a normal person would refuse.
The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
Semantic web was a separate initiative by Tim Berners-Lee when the web already largely used un-semantic HTML. And it never went anywhere.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Well they had an API, but…
Reddit does provide RSS feeds, e.g.: https://www.reddit.com/r/SonicTheHedgehog/.rss
Frankly, I’m surprised they still offer RSS feeds. They’ve been slowly but surely killing off all ways of accessing their content for years. One day they’ll disable them, but for now they still work.
so why was I getting hit with over 1,400,000 request a day to the web URI and not the API by some bot farm in China the other week. They were also hitting other lemmy instances.
I blocked the fuckers, no qualms at all.
Even if they were using the API they were not being nice about their shit.
Reddit does have RSS feeds
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
deleted by creator
Its amazing to me how consistently the shitty behaviours of these billionaire techbro oligarchs impact disabled or marginalised people… even when the point isnt to directly shit on them. Its fucking vile.
I honestly think many (too many, but certainly not all! I am one) programmers are some of the immoral, ethically spurious people around in the 21st century.
That is why they want AI to replace programmers. AI morals are programmed, so they can be designed to do shitty things that a normal person would refuse.
Hmffh, anti-copyright. After all, every view is a copy to your machine. Just information being free.