【文章标题】:从互联网博客中全选(Select * from Internet.blogposts)

【文章正文】: In 2007, Tim Berners Lee write the essay The Giant Global Graph. Quote: 2007年,蒂姆·伯纳斯·李撰写了《巨型全球图谱》一文。文中引述: “There are cries from the heart .. for my friendship, that relationship to another person, to transcend documents and sites. ..Then any other site or program can use that information.” “人们发自内心地渴望…让友谊这种人际关系能超越文档和网站…这样任何其他网站或程序都能使用这些信息”

It seems appropriate as X is sending cease-and-desist letters to Nitter to remember TBL’s essay. Nitter is - was - a simple frontend to X which allows users to view tweets without logging in. Even that small use of proxying to the pages is enough to receive threats of legal action. 当X平台向Nitter发送禁止函时,重读伯纳斯·李的论述显得尤为应景。Nitter曾是X平台的简易前端,允许用户无需登录即可浏览推文。即便是这种简单的页面代理使用,也足以招致法律诉讼威胁。

Twitter’s API in 2007 was famously open, which meant thousands of developers building clients, tools, and analytics for free. So, what happened? Why was it pulled? Simple: the network won. The developers stopped being an asset, and the API progressively closed. Rate limits, pricing tiers, login requirements, then technical blocks on the workarounds, and now letters from lawyers. Meta ran the same playbook a decade ago and it’s now hard to remember there was ever a Facebook or Instagram API worth building on. 2007年推特API以开放著称,数千开发者免费构建客户端、工具和分析系统。后来发生了什么?为何开放政策被取消?很简单:网络效应赢了。开发者从资产变成负担,API逐渐封闭:先是速率限制、分级收费、登录要求,再到技术封堵变通方案,如今是律师函警告。Meta十年前就采用相同策略,如今已难想象Facebook或Instagram曾有过值得构建的API。

This is why Brewster Kahle, founder of the Internet Archive, has been calling for over a decade for us to lock the Web open. 这正是互联网档案馆创始人布鲁斯特·卡勒十余年来呼吁”锁定网络开放”的原因。

Nitter started off using X’s APIs. When that closed, it read public web pages. And now that there’s nothing left to close, the demand is that the source code come down. A program that displays public posts is being treated as a circumvention device under computer-crime statutes. Nitter最初使用X的API,当API关闭后转为抓取公开网页。如今再无封锁余地时,对方竟要求下架源代码。一个展示公开内容的程序竟被视作计算机犯罪法规下的规避工具。

We have a walled garden problem. It isn’t going to change, and the only option in front of us is to start fresh. 我们面临围墙花园困境。现状不会改变,唯一选择是推倒重来。

The good news is, atproto continues to grow, activitypub remains resilient, and our community is full of believers and builders in the open social web. Since I work on atproto, that’s what I’ll talk about next. 好消息是atproto持续发展,ActivityPub保持韧性,开放社交网络社区充满信徒与建设者。鉴于我从事atproto开发,接下来将重点讨论它。

Interoperation by SELECT * 通过SELECT *实现互操作 SELECT * FROM internet.blogposts 从互联网博客中全选

The walled garden problem is downstream of a simple question: how do I SELECT * FROM internet? 围墙花园问题源于一个简单疑问:如何从互联网全选数据?

If you’ve never written database code, SELECT * FROM users is how you ask a database for everything it knows about its users. Once you have it you can filter it, sort it, and join it against anything else you’ve got. 若非程序员,SELECT * FROM users即向数据库索取全部用户数据。获得数据后即可过滤、排序、关联其他数据集。

The web doesn’t historically work that way. The web is a few dozen companies, each holding a filing cabinet, each with a receptionist posted out front. He’ll read you one file at a time, but only files you can name, as fast as he cares to read, and as long as his boss allows. 但互联网从未如此运作。网络由数十家公司掌控,每家如同配有前台接待的档案柜。接待员每次只按老板规定速度、仅读取你能指名的文件。

Nitter was a lightweight X reader that worked fine right up until X turned off the access it depended on. Every API (the “receptionist”) is a business decision that hasn’t been reversed yet. Nitter作为轻量级X阅读器,在X切断其访问权限前一直运作良好。每个API(即”接待员”)都是尚未撤销的商业决策。

But Impermanence isn’t the only problem. Even a permanent, free, generously rate-limited API wouldn’t be enough. Applications need much more meaningful access than APIs can provide. 但无常性并非唯一问题。即使永久免费且宽松限速的API也不够。应用需要比API更本质的访问权限:

  • You can only ask questions someone already thought to answer. An API is a fixed menu. It gives you getPosts(user) andgetFollowers(user) . If your product idea needs “posts from people my followers follow, ranked by how often they get quoted,” there is no endpoint for that, and there never will be, because nobody at that company is building for your product.

  • 你只能提出预设问题。API是固定菜单,提供getPosts(用户)和getFollowers(用户)。若你需要”按被引用频率排序的关注者所关注人的帖子”,没有现成接口且永远不会有,因该公司无人为你的产品开发。

  • Even the right questions come back in the wrong shape. Followers come 100 at a time. A two-million-follower account is 20,000 round trips. At any polite rate limit that’s hours of work to answer one question about one user — so anything interactive, anything that has to feel instant, is off the table before you start.

  • 正确问题也返回错误形式。关注者每次返回100条。200万粉账号需2万次请求。任何礼貌的速率限制下,回答单个用户问题都需数小时——任何需要即时感的交互功能在开始前就已出局。

  • You can’t join across “cabinets”. The interesting questions are almost always cross-service: this person’s posts against that person’s photos against a third service’s reviews. Two receptionists in two buildings can’t cross-reference anything, and neither can you.

  • 无法跨”档案柜”关联。有趣问题常需跨服务:某人的帖子关联另一人的照片再关联第三方评论。两栋楼里的接待员无法交叉引用,你也一样。

  • You can’t index data you don’t hold. Search, ranking, recommendations, feeds, moderation tooling — all of it is built on indexes over the whole corpus, laid out for the specific questions your product asks. You cannot build an index through a keyhole.

  • 无法索引非持有数据。搜索、排序、推荐、订阅流、审核工具——全都依赖全局语料库索引,按产品需求定制。你无法通过锁眼构建索引。

To actually build a service, we need the whole dataset rather than a view onto it; we need it live, arriving as it changes instead of polled for; we need to index it however my product demands; we need to write back into it; and we need all of that guaranteed in a way no single company’s quarterly priorities can revoke. 真正构建服务需要完整数据集而非视图;需要实时推送而非轮询;需按产品需求自由索引;需写入权限;且需确保这些权利不因某公司季度目标而撤销。

Desktop apps handle this by sharing the filesystem. Internet apps don’t use files; they use databases. We need to share the database. 桌面应用通过共享文件系统实现。互联网应用不用文件而用数据库。我们需要共享数据库。

As a user, I don’t want to be locked into an app anymore than I’d want to be locked in the trunk of a car. I want an actual free market. 作为用户,我不想被锁在应用中,就像不想被锁进汽车后备箱。我要真正的自由市场。

So then, here’s another set of needs. 因此需要另一组条件:

  • Persistence of identity.

    • My presence and relationships are built around my identity. It needs to outlive the app I signed up with.
  • 身份持久性

    • 我的存在和社交关系围绕身份建立。身份必须比注册应用更长寿。
  • The export of living (not dead) data between services.

    • Exporting archives of your tweets is useless as an account migration solution because data doesn’t live in isola
  • 服务间活跃数据(非死数据)导出

    • 导出推文存档对账户迁移无意义,因数据无法在孤立状态下存活