Developer Creates Infinite Maze That Traps AI Training Bots

kororon@lemmy.cafe · 1 year ago

Developer Creates Infinite Maze That Traps AI Training Bots

daniskarma@lemmy.dbzer0.com · 1 year ago

Yeah, that has like 0 chances for working. At most it would annoy bots for web search, at least it has a proper robots.txt.

But any agent trying to process data for AI is not going to go to random websites. It’s going to use a curated list of sites with valuable content.

At this point text generation datasets can be achieved with open data, and data sold by companies like reddit or Microsoft, they don’t need to “pirate” your blog posts.

ShortFuse@lemmy.world · 1 year ago

scrape.maxDepth = 5

Lovable Sidekick@lemmy.world · 1 year ago

LOL wow, this is probably the most elegant way to say what I just said to somebody else. Well written web crawlers aren’t like sci-fi robots that rock back and forth smoking when they hear something illogical.

brb@sh.itjust.works · 1 year ago

What’s stopping the sites with valuable content from using this?

FaceDeer@fedia.io · 1 year ago

A bot that’s ignoring robots.txt is likely going to be pretending to be human. If your site has valuable content that you want to show to humans, how do you distinguish them from the bots?

notastatist@feddit.org · 1 year ago

What is robots.txt?

Traister101@lemmy.today · 1 year ago

A file that “robots” are supposed to respect when they index a website. Here’s Googles https://www.google.com/robots.txt

FaceDeer@fedia.io · 1 year ago

A standard method for indicating to web-crawlers what pages they should read and which they should skip.

nucleative@lemmy.world · 1 year ago

I think sites that feel they have valuable content can deploy this and hope to trap and perhaps detect those bots based on how they interact with the tarpit

Lovable Sidekick@lemmy.world · 1 year ago

True to a limited extent. Anyone can post a link to somebody’s blog on a site like reddit without the blogger’s permission, where a web crawler scanning through posts and comments would find it. But I agree with you that a thing like Nepehthes probably wouldn’t work. Infinite loop detection is an important part of many types of software and there are well-known techniques for it, which as a developer I would assume a well written AI web crawler would have (although I’ve never personally made one).