【文章标题】:Building an AI Text Detector From Scratch

【文章标题中文翻译】:从头构建一个AI文本检测器

【文章正文】: Substack recently launched its AI detector feature in the UI, which is super interesting.

【中文翻译】: Substack最近在其界面中推出了AI检测功能,这非常有趣。

Separately, lots of people asked me about interesting local do-it-yourself LLM projects as demos to show what small language models (SLMs) are capable of.

【中文翻译】: 另外,很多人问我关于有趣的本地自制LLM项目,作为演示小语言模型(SLM)能力的示例。

Putting one and one together, I thought it would be interesting to show how an AI detector can be implemented. I will also use it as a verifier to train a small language model to produce text that avoids detection. This is a small educational project for studying the limitations of AI detectors and exploring a verifier-based LLM application beyond regular reasoning models trained on math and code.

【中文翻译】: 把这两件事结合起来,我觉得展示如何实现一个AI检测器会很有意思。我还会把它用作验证器,训练一个小语言模型生成能规避检测的文本。这是一个小型教育项目,旨在研究AI检测器的局限性,并探索基于验证器的LLM应用,超越那些在数学和代码上训练的常规推理模型。

Figure 1: Substack now features a built-in AI detector.

【中文翻译】: 图1:Substack现在内置了AI检测功能。

So, as mentioned above, the intended goal of this tutorial is to explain how AI detectors work by building (a simple) one.

【中文翻译】: 因此,如上所述,本教程的目标是通过构建一个(简单的)检测器来解释AI检测器的工作原理。

In practice, such a detector can be used to filter out spammy content, but also to potentially improve your personal writing without turning it into AI-generated text. For example, if you wrote a lengthy article and want to improve spelling and grammar, it is tempting (and actually useful) to use a grammar checker to polish it and improve readability. There are different services for that, including general-purpose LLMs like ChatGPT. However, this also runs the risk that these tools turn your writing, even though it’s still your own writing, into something that is then overpolished and now sounds like AI and gets flagged as spammy content.

【中文翻译】: 在实践中,这样的检测器可以用来过滤垃圾内容,也可能在不把你的个人写作变成AI生成文本的前提下改进你的写作。例如,如果你写了一篇长文章,想要改进拼写和语法,使用语法检查器来润色并提高可读性是很有吸引力的(也确实有用)。有不同的服务可以做到这一点,包括像ChatGPT这样的通用LLM。然而,这也有风险:这些工具可能会把你的写作——尽管仍然是你自己的写作——变成过度润色、听起来像AI的内容,并被标记为垃圾内容。

For example, with an AI checker, one could say, “Fix my grammar while ensuring that my text still scores 0% AI-generated.”

【中文翻译】: 例如,使用AI检查器,人们可以说:“修复我的语法,同时确保我的文本仍然保持0%的AI生成评分。”

Anyway, while we are building a fully functional checker here, the goal is to explain 1) how AI checkers (can) work and 2) use this as a case study for a more general topic on how to build a scorer or verifier that can be used with LLMs.

【中文翻译】: 无论如何,虽然我们在这里构建的是一个功能完整的检查器,但目标是解释1)AI检查器(可以)如何工作,以及2)将其作为更通用主题的案例研究,即如何构建可用于LLM的评分器或验证器。

Disclaimer: AI checkers are essentially a cat-and-mouse game. AI checkers may learn to detect a certain pattern that is indicative of AI-generated content. Then, the next LLM may incidentally or deliberately not exhibit that pattern and avoid detection. The AI checker then has to be updated to detect said LLM, and so forth. Plus, it’s also likely to encounter false positives (human written text flagged as AI-generated), but more on that later.

【中文翻译】: 免责声明:AI检测器本质上是一场猫鼠游戏。AI检测器可能会学会检测某些表明AI生成内容的模式。然后,下一代LLM可能偶然或故意不表现出这种模式,从而避免被检测到。AI检测器随后必须更新以检测该LLM,如此循环。此外,它还可能遇到误报(人类书写的文本被标记为AI生成),但稍后会详细讨论。

Project goals

【中文翻译】: 项目目标

There are several goals of this project. The overarching goal is, of course, to illustrate how AI detectors work and show an applied end-to-end LLM project including evaluation, training, and local deployment for real-world use.

【中文翻译】: 这个项目有几个目标。首要目标当然是说明AI检测器的工作原理,并展示一个实际应用的端到端LLM项目,包括评估、训练和本地部署,以供实际使用。

The outcome of this is an AI-detector API that can be used by humans and agents, and a user-friendly UI.

【中文翻译】: 最终成果是一个可供人类和智能体使用的AI检测器API,以及一个用户友好的界面。

Figure 2: Preview of the local browser interface developed later in this project. It returns a whole-text AI score and can also highlight the scores for individual text chunks.

【中文翻译】: 图2:本项目稍后开发的本地浏览器界面预览。它返回整段文本的AI评分,并可以高亮显示各个文本块的评分。

Method overview

【中文翻译】: 方法概述

Here, we are going to develop a method similar to Pangram models, which, as far as I know, are behind Substack AI detection feature.

【中文翻译】: 在这里,我们将开发一种类似于Pangram模型的方法,据我所知,Substack的AI检测功能背后使用的就是这种模型。

I wrote a short article about AI-text detection a while back in 2023:

【中文翻译】: 我在2023年早些时候写过一篇关于AI文本检测的短文:

What Are the Different Approaches for Detecting Content Generated by LLMs Such As ChatGPT? And How Do They Work and Differ?

【中文翻译】: 有哪些不同的方法来检测由ChatGPT等LLM生成的内容?它们如何工作,又有何不同?

In essence, there are different ways to detect AI-written text, from supervised classifiers and perturbation-based probability tests to perplexity measures and watermarking.

【中文翻译】: 本质上,检测AI书写文本有多种方式,从监督分类器和基于扰动的概率测试,到困惑度度量和水印技术。

In this tutorial, we will build a model that returns a 0-100 score. It’s essentially a classifier with an estimated probability score. The probability score will denote how likely a text is AI-generated according to the classifier. (Or, to be precise the score is the classifier’s estimated probability for the AI-generated class based on its training distribution. However, we shouldn’t interpreted it as a general probability that the text was written by AI.)

【中文翻译】: 在本教程中,我们将构建一个返回0-100分的模型。它本质上是一个带有估计概率分数的分类器。该概率分数表示根据分类器,一段文本是AI生成的可能性有多大。(或者,准确地说,该分数是分类器基于其训练分布对AI生成类别的估计概率。然而,我们不应将其解释为文本由AI撰写的一般概率。)

For this, we are going to fine-tune a DistilBERT classifier (similar to what I described in one of my early Substack articles,

Finetuning Large Language Models

), but more details on that later when we get to that stage.

【中文翻译】: 为此,我们将微调一个DistilBERT分类器(类似于我在早期Substack文章中描述的内容:

微调大型语言模型

),但更多细节将在我们进入该阶段时再讨论。

Read more

【中文翻译】: 阅读更多