【文章标题】:Creepy crawlies
【文章标题】:令人毛骨悚然的网络爬虫
【文章正文】:
Creepy crawlies
令人毛骨悚然的网络爬虫
Konstantin Ryabitsev discusses how bad the “background radiation” of abusive crawlers has become from the perspective of
Konstantin Ryabitsev从Linux内核官方Git仓库git.kernel.org的视角,讨论了恶意爬虫造成的”背景辐射”问题已严重到何种程度:
git.kernel.org
TL;DR: we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones. At any one time, across 5 geo-distributed nodes, there are 14 CPU cores doing nothing but rendering git commits as html.
简而言之:我们为爬虫渲染提交记录所消耗的CPU周期,已超过包括git克隆在内的所有合法访问的总和。在任意时刻,跨越5个地理分布式节点,都有14个CPU核心专门用于将git提交记录渲染成html。
I worry about this a lot from the perspective of Datasette, which serves a huge number of crawlable web pages.
从Datasette(提供大量可爬取网页服务)的视角来看,我对此深感忧虑。
Via
来源:
Hacker News
黑客新闻
Tags:
标签:
crawling
网络爬虫
,
、
git
git版本控制系统
,
、
linux
Linux操作系统
,
、
datasette
Datasette数据库工具
,
、
ai-ethics
人工智能伦理