【文章标题】:Creepy crawlies
【文章标题】:令人毛骨悚然的网络爬虫

【文章正文】:
Creepy crawlies
令人毛骨悚然的网络爬虫

Konstantin Ryabitsev discusses how bad the “background radiation” of abusive crawlers has become from the perspective of
Konstantin Ryabitsev从Linux内核官方Git仓库git.kernel.org的视角,讨论了恶意爬虫造成的”背景辐射”问题已严重到何种程度:

git.kernel.org

TL;DR: we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones. At any one time, across 5 geo-distributed nodes, there are 14 CPU cores doing nothing but rendering git commits as html.
简而言之:我们为爬虫渲染提交记录所消耗的CPU周期,已超过包括git克隆在内的所有合法访问的总和。在任意时刻,跨越5个地理分布式节点,都有14个CPU核心专门用于将git提交记录渲染成html。

I worry about this a lot from the perspective of Datasette, which serves a huge number of crawlable web pages.
从Datasette(提供大量可爬取网页服务)的视角来看,我对此深感忧虑。

Via
来源:

Hacker News
黑客新闻

Tags:
标签:

crawling
网络爬虫

,
、

git
git版本控制系统

,
、

linux
Linux操作系统

,
、

datasette
Datasette数据库工具

,
、

ai-ethics
人工智能伦理