【文章标题】:Bot Detection Without JavaScript: What My Blog Measured
【无需JavaScript的机器人检测:我的博客实测数据】
【文章正文】:
Bot Detection Without JavaScript: What My Blog Measured
【无需JavaScript的机器人检测:我的博客实测数据】
On my blog, network and request-header rules moved 277 of 372 browser-User-Agent requests out of the Browsers category: 74.5%. That gives me a much more useful account of the traffic arriving at my Cloudflare Worker. It has not established how many people read the site. Over the same two complete UTC days, the remaining 95 Browser HTML observations still differ from 14 Cloudflare Web Analytics page loads. The useful result is knowing which requests the rules separate, why they separate them, and where the evidence stops.
【在我的博客上,网络和请求头规则将372个浏览器User-Agent请求中的277个移出了“浏览器”类别(占比74.5%)。这使得我能更准确地分析到达Cloudflare Worker的流量。但这并未确认实际阅读人数。在相同的两个完整UTC日内,剩余的95次浏览器HTML访问记录仍与Cloudflare网络分析的14次页面加载量存在差异。关键价值在于明确规则如何筛选请求、筛选依据及证据链的边界。】
-
Network evidence catches requests that pass the header checks. In that window, 60 cloud-classified requests carried the navigation headers the browser rule requires.
【- 网络证据会捕获通过头部检查的请求。该时段内有60个云端分类请求携带了浏览器规则要求的导航标头。】 -
A reason for each classification makes the counter explainable. It also exposed mistakes: our HTML-acceptance check mishandled valid headers, and new rows were missing their network-provenance marker. Both have since been repaired.
【- 每个分类理由使计数器具有可解释性。同时也暴露了错误:我们的HTML接受检查误判了有效标头,且新增数据缺失网络来源标记。这两点均已修复。】 -
Client identity and readership need different evidence. Nine stored signature verifications identify signers, including crawlers and deliberate tests. They do not count people asking an assistant to read.
【- 客户端身份与读者量需不同证据。9条存储的签名验证可识别请求方(包括爬虫和主动测试),但不统计通过语音助手访问的用户。】 -
The remaining disagreement is a measured problem. Neither a smaller Browser count nor agreement with a script counter establishes audience accuracy.
【- 剩余差异是经测量的客观问题。无论浏览器计数偏低还是与脚本计数器一致,均不能证明受众数据的准确性。】
Comparing edge page views with a script counter
【边缘页面浏览量 vs 脚本计数器对比】
Comparing two counters exposed the problem with my first-party analytics: I had treated browser User-Agents as evidence of readers. The counters measure different events, so their disagreement is a starting point for investigation. It cannot, by itself, tell me which requests were automation.
【对比两个计数器暴露了第一方分析工具的问题:我曾将浏览器User-Agent视作读者证据。由于二者测量不同事件,其差异正是调查的起点。但仅凭差异无法判定自动化请求。】
The initial alarm came from this comparison, saved on September 3:
【9月3日的对比数据触发初步警报:】
| Source | Page events / loads | Client or visit metric |
|---|---|---|
| D1 browser-UA class, seven UTC days ending September 2 | 1,209 | 578 daily client identifiers |
| Cloudflare Web Analytics, its rolling seven-day dashboard window | 113 | 52 visits |
The 578-versus-52 difference looked like eleven times as many readers. But a daily client identifier is not a visit, the time windows were not identical, and the script dashboard included /stats. Even the more comparable 1,209-versus-113 page totals needed those qualifications. The original query record preserves the comparison as it was made.
【578 vs 52的差异看似读者量相差11倍。但每日客户端标识≠访问次数,时间窗口不同,且脚本统计包含/stats路径。就连更可比的1,209 vs 113页面总量也需考虑这些前提。原始查询记录保留了当时的对比状态。】
There was stronger evidence inside the requests. On September 2, 100 of 113 daily clients loaded one page; 156 of 164 browser-UA page observations carried no referrer. Those facts alone would not establish automation. One client classified as mobile, however, fetched 31 distinct pages in the same timestamp second. My note that night was:
【请求内部存在更强证据:9月2日,113个每日客户端中有100个仅加载单页;164次浏览器-UA页面访问中有156次无来源页。这些单独来看不能证明自动化。但一个标记为移动端的客户端在同一秒请求了31个不同页面。当晚我的笔记如下:】
i am seeing daily clients as 113 for today, and it seems unbelievable to me, like which articles are they reading, where are they coming from and so on… i just published a new article and its not even coming up in the Top pages by views section… like whats going on… Like sometimes when i publish the post i wanna see for this particular post how many readers have arrived and through which sources, its impossible to figure out. But still the most bizarre is the numbers, who is all reading these articles, it seems insane, which i appreciate but I don’t want to gaslight ourselves, like something is not adding up
【今日显示113个独立客户端,难以置信——他们在读哪些文章?来自何处?…刚发布的新文章甚至没进入浏览量Top榜…发布后想查看特定文章的读者来源时完全无法追踪…最诡异的是这些数字,究竟谁在读?数据明显异常,虽受宠若惊但必须保持清醒——某些环节肯定有问题】
A useful comparison needs four decisions made before calculating the ratio:
【有效的对比需在计算比率前明确四点:】
-
Choose the same host and explicit start-inclusive, end-exclusive UTC window.
【- 选择相同主机和明确的UTC时间窗口(含起始,不含结束)】 -
Compare page events with page loads. Report daily identifiers and visits separately.
【- 对比页面事件与页面加载量。分别报告每日标识符和访问次数】 -
Record which routes, response types, bots, owners, and test requests each system excludes. Keep sampling information with the result.
【- 记录各系统排除的路由、响应类型、机器人、所有者和测试请求。结果需附带抽样信息】 -
Inspect the discrepant requests and test the possible collection differences.
【- 检查差异请求并测试可能的收集偏差】
Cloudflare documents script blockers and browser or network loss as reasons its beacon can miss page loads. Browser caching and differing eligibility also need checking against the edge counter. A script can run in an automated browser. There is no universal ratio that separates these causes. Cloudflare Web Analytics FAQ, checked September 6, 2026.
【Cloudflare官方文档指出脚本拦截器、浏览器/网络故障会导致信标遗漏页面加载。浏览器缓存和不同准入标准也需与边缘计数器核对。脚本可在自动化浏览器运行。不存在区分这些原因的通用比率。——2026年9月6日查阅的Cloudflare分析FAQ】
The comparison reveals questions we can investigate. The earlier version of this article’s under-two validation threshold was unsupported, and I have removed it.
【对比揭示了可调查的问题。本文早期版本中“低于2的验证阈值”缺乏依据,现已删除。】
The request rules in the Cloudflare Worker
【Cloudflare Worker中的请求规则】
The Worker classifies recorded request characteristics. It can apply those rules deterministically without establishing who controlled the client. This section is for someone implementing the classifier; the measured results below can be read without the implementation detail.
【Worker根据记录的请求特征进行分类。它可确定性地应用规则而无需确认客户端控制者。本节面向分类器实施者,下方测量结果可不关注实现细节。】
The edge counter schedules a D1 observation after an eligible successful page GET. It excludes prefetches, /stats, API routes, and non-page responses. It includes HTML and negotiated Markdown page responses; direct .md requests are outside this counter. That collection boundary comes from the eligibility code.
【边缘计数器在符合条件的成功GET请求后安排D1观测。排除预加载、/stats、API路由和非页面响应。包含HTML和协商式Markdown页面响应(直接.md请求不统计)。该收集边界源自准入代码。】
Four
【四】(注:此处原文截断)