Creepy Crawlies(people.kernel.org)
1333 points by zdw 11 days ago | 679 comments
tl;dr: AI scrapers are hammering git.kernel.org by crawling every commit URL across 922 forks of linux.git as HTML instead of just cloning the repos, consuming ~20% of total CPU capacity across 5 nodes purely to render commits for bots. Defenses like IP/ASN bans failed once crawlers shifted to residential proxy networks, and Anubis proof-of-work challenges (now at difficulty 5) are being solved by ~33% of bots. Legitimate traffic is estimated at just 2%, forcing kernel.org to disable features and gate expensive operations.
HN Discussion:
  • Anubis proof-of-work is fundamentally flawed since scrapers handle it better than mobile users
  • ~Alternative defenses like tarpits, traps, or obscurity forks would be more effective than Anubis
  • Crawlers are indiscriminate and thoughtless, explaining why even niche cgit instances get hammered
  • Sharing personal war stories of being overwhelmed by bots and forced to disable features
  • AI companies themselves incentivize disregard for crawling etiquette like robots.txt