【文章标题】:When str.lower() is a security vulnerability in Python – Seth Larson 【文章标题】:当 str.lower() 成为 Python 中的安全漏洞 – Seth Larson

Some internet standards only support ASCII characters, but the world uses much more than the Latin alphabet. Thus, a mapping from Unicode to ASCII for use in domain names is required. 一些互联网标准仅支持 ASCII 字符,但世界使用的字符远不止拉丁字母。因此,需要一种从 Unicode 到 ASCII 的映射,用于域名。

NamePrep was part of that solution, defined in RFC 3491 as a profile of StringPrep, and is crucially a component of Internationalizing Domain Names in Applications (IDNA), also known as “IDNA 2003”. The StringPrep algorithm is defined in RFC 3454. IDNA 2003 has been obsoleted by IDNA 2008 defined in RFC 5890, 5891, 5892, and 5893. NamePrep 是解决方案的一部分,在 RFC 3491 中被定义为 StringPrep 的一个配置文件,并且是应用程序国际化域名(IDNA,也称为“IDNA 2003”)的关键组成部分。StringPrep 算法在 RFC 3454 中定义。IDNA 2003 已被 RFC 5890、5891、5892 和 5893 中定义的 IDNA 2008 取代。

Python supports IDNA 2003 through the idna codec (str.encode(‘idna’)) and IDNA 2008 is supported by the idna package on the Python package Index. Python’s implementation of StringPrep is implemented in the stringprep module in the standard library. In general, you should be using the idna package (IDNA 2008) and not .encode(“idna”) (IDNA 2003), but sometimes you do need the older behavior. Python 通过 idna 编解码器(str.encode(‘idna’))支持 IDNA 2003,而 IDNA 2008 由 Python 包索引上的 idna 包支持。Python 的 StringPrep 实现位于标准库的 stringprep 模块中。一般来说,你应该使用 idna 包(IDNA 2008),而不是 .encode(“idna”)(IDNA 2003),但有时你确实需要旧行为。

StringPrep defines the “case folding” step (case folding is approximately “how to lowercase/uppercase a codepoint”) in Section 3.2, enabling case-insensitive comparisons of strings, by mapping all characters through mapping tables B.2 and B.3. B.2 is effectively str.lower(), lowercasing all characters according to Unicode rules and B.3 contains the exceptions. The Python code implementing this (and assuming B.3 table is captured correctly) is the following code below: StringPrep 在第 3.2 节中定义了“大小写折叠”步骤(大小写折叠大致是“如何将码点转换为小写/大写”),通过映射表 B.2 和 B.3 映射所有字符,实现字符串的大小写不敏感比较。B.2 实际上就是 str.lower(),根据 Unicode 规则将所有字符转换为小写,而 B.3 包含例外情况。实现此功能的 Python 代码(假设 B.3 表已正确捕获)如下所示:

def map_table_b3(code): r = b3_exceptions.get(ord(code)) if r is not None: return r return code.lower() And that might seem fine… and the title probably gave it away already. 这看起来可能没问题……而标题可能已经透露了答案。

The str.lower() call in this function is a vulnerability! 这个函数中的 str.lower() 调用是一个漏洞!

Why? Because str uses whatever Unicode data that the particular Python interpreter is shipped with, you can figure out what Unicode version your Python interpreter uses by accessing unicodedata.unidata_version: 为什么?因为 str 使用特定 Python 解释器自带的 Unicode 数据,你可以通过访问 unicodedata.unidata_version 来查明你的 Python 解释器使用的 Unicode 版本:

import unicodedata unicodedata.unidata_version ‘17.0.0’ There’s also a database of Unicode 3.2.0 data available on every version of Python (unicodedata.ucd_3_2_0) specifically for the StringPrep and IDNA algorithms: 此外,每个版本的 Python 都提供了 Unicode 3.2.0 数据数据库(unicodedata.ucd_3_2_0),专门用于 StringPrep 和 IDNA 算法:

$ grep -I “ucd_3_2_0” -R Lib/ Lib/stringprep.py:from unicodedata import ucd_3_2_0 as unicodedata Lib/encodings/idna.py:from unicodedata import ucd_3_2_0 as unicodedata This is important! StringPrep depends on this specific version of Unicode to operate consistently, the B.2 and B.3 tables in RFC 3454 are essentially Unicode 3.2.0 case-folding rules encoded into a table. So we need to use Unicode 3.2.0 case-folding rules, not newer Unicode case-folding rules. This is why calling str.lower() represents a difference in the implementation and the specification, and therefore a vulnerability: 这很重要!StringPrep 依赖这个特定版本的 Unicode 才能一致运行,RFC 3454 中的 B.2 和 B.3 表本质上就是编码到表中的 Unicode 3.2.0 大小写折叠规则。因此我们需要使用 Unicode 3.2.0 的大小写折叠规则,而不是更新的 Unicode 大小写折叠规则。这就是为什么调用 str.lower() 代表了实现与规范之间的差异,因此是一个漏洞:

RFC 3454 compliant value (‘Ꭰ’ is U+13A0)

“ᎠᎠ”.encode(“idna”) ‘xn—58da’

Value if using Unicode 17.0.0 case-folding

“ᎠᎠ”.encode(“idna”) ‘xn—kz9aa’ The fix was to create new exceptions so that str.lower() would behave as if it was using Unicode 3.2.0 for only particular function. So, we go through each Unicode codepoint and record when the behavior of str.lower() is different when comparing the Unicode version shipped with Python and Unicode 3.2.0. 修复方法是创建新的例外,使 str.lower() 在特定函数中表现得如同使用 Unicode 3.2.0。因此,我们遍历每个 Unicode 码点,记录当比较 Python 自带的 Unicode 版本与 Unicode 3.2.0 时 str.lower() 的行为差异。

And that’s all, now IDNA 2003 is consistent with the specification. 就是这样,现在 IDNA 2003 与规范一致了。

Thanks to Bitshift for reporting the vulnerability, Stan Ulbrych for co-developing the remediation, and Marc-Andre Lemburg and Petr Viktorin for reviewing the remediation. See CVE-2026-17084 for more details. 感谢 Bitshift 报告此漏洞,Stan Ulbrych 共同开发修复方案,以及 Marc-Andre Lemburg 和 Petr Viktorin 审查修复方案。有关更多详细信息,请参阅 CVE-2026-17084。

My work as the Security Developer-in-Residence at the Python Software Foundation is sponsored by Alpha-Omega. Thanks to Alpha-Omega for supporting security in the Python ecosystem. 我作为 Python 软件基金会驻场安全开发人员的工作由 Alpha-Omega 赞助。感谢 Alpha-Omega 对 Python 生态系统安全的支持。

    Wow, you made it to the end!
    哇,你读到了最后!
  • Share your thoughts with me on Mastodon, email, or Bluesky.

  • 在 Mastodon、电子邮件或 Bluesky 上与我分享你的想法。

  • Browse this blog’s archive of 193 entries.

  • 浏览本博客的 193 篇文章存档。

  • Check out this list of cool stuff I found on the internet.

  • 查看我在互联网上发现的有趣内容列表。

  • Follow this blog on RSS or the email newsletter.

  • 通过 RSS 或电子邮件通讯关注本博客。

  • Go outside (best option)

  • 出去走走(最佳选择)