【文章标题】:LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
【标题翻译】:大语言模型评审员验证存在而非缺失:AI临床笔记中的遗漏盲区及其破解之道
【文章正文】:
Computer Science > Computation and Language
[Submitted on 31 Aug 2026]
Title:LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It
计算机科学 > 计算与语言
[2026年8月31日提交]
标题:大语言模型评审员验证存在而非缺失:AI临床笔记中的遗漏盲区及其破解方法
Abstract:Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant.
摘要:环境AI助手起草临床笔记,已发布的审计报告显示其主要错误是遗漏:诊疗过程中确认但未被记录的信息。标准检查采用大语言模型评审员:由第二个模型对照诊疗记录审阅笔记并标记问题。我们探究评审员能否检测遗漏。公开语料库无法提供答案:其临床医生参考笔记与原始记录存在实质性差异。
Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions.
我们的基准测试包含500对来自审计事实清单的单错误笔记,其中298例明确缺失特定事实,202例为添加/篡改的对照组。在八种评审设计中,配对判别(有缺陷笔记低于其纯净版本,0.5为随机概率)对添加/篡改内容的识别率为0.79-0.94,而对遗漏的识别率仅为0.50-0.63。
On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each.
在单笔记评审中,所有设计对遗漏的标记频率均未显著高于完美笔记。措辞调整、投票机制和GEPA提示优化虽能改变操作点,但未形成有效检测。重构任务可破解该问题:先列出诊疗记录确认的事实,再逐项核对笔记。
Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline’s flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note.
两种方法独立实现并形成权衡:逐事实处理流程 vs GEPA进化提示的单次调用。流程法能以2.7%误报率标注缺失事实及其严重程度;单次调用法检测率更高(36.9%对24.6%,p=0.002),误报率6.2%且单笔记成本降低90%。
A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate.
医师作者验证70个项目,在两种方法分歧时10/10次支持流程法(p=0.002)。非作者的第二位临床医生盲评严重程度量表,结果误差不超过一级。在实际供应商笔记测试中,基准阈值无法直接迁移,但重新校准的单次调用法以半数误报率超越八种设计中的最佳表现。
Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.
事实在其他位置重述的遗漏会同时突破两种方法。我们公开基准测试集、提示模板及判断结果。
References & Citations
参考文献与引用
Loading…
载入中…
Bibliographic and Citation Tools
文献计量与引用工具
Bibliographic Explorer (What is the Explorer?)
文献浏览器(什么是浏览器?)
Connected Papers (What is Connected Papers?)
关联论文(什么是关联论文?)
Litmaps (What is Litmaps?)
文献图谱(什么是文献图谱?)
scite Smart Citations (What are Smart Citations?)
scite智能引用(什么是智能引用?)
Code, Data and Media Associated with this Article
本文相关代码、数据与媒体
alphaXiv (What is alphaXiv?)
alphaXiv(什么是alphaXiv?)
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
CatalyzeX论文代码检索(什么是CatalyzeX?)
DagsHub (What is DagsHub?)
DagsHub(什么是DagsHub?)
Gotit.pub (What is GotitPub?)
Gotit.pub(什么是GotitPub?)
Hugging Face (What is Huggingface?)
Hugging Face(什么是Hugging Face?)
ScienceCast (What is ScienceCast?)
ScienceCast(什么是ScienceCast?)
Demos
演示
Recommenders and Search Tools
推荐与检索工具
Influence Flower (What are Influence Flowers?)
影响力花图(什么是影响力花图?)
CORE Recommender (What is CORE?)
CORE推荐系统(什么是CORE?)
arXivLabs: experimental projects with community collaborators
arXivLabs:与社区合作者共同开展的实验项目
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
arXivLabs是一个允许合作者直接在我们网站上开发和共享新功能的框架。
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
与arXivLabs合作的个人和组织都认同我们关于开放、社区、卓越和用户数据隐私的价值观。arXiv坚持这些价值观,仅与遵守这些价值观的伙伴合作。
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
有能为arXiv社区创造价值的项目构想?了解更多关于arXivLabs的信息。
🔗 知识库双向关联
- [[Raw_翻译_[AINews] Death of Params Z.ai CEO Jie Tang on GLM 5.3 and th|Death of Params Z.ai CEO Jie Tang on GLM 5.3 and th_全翻译]]
- Building an AI Text Detector From Scratch_全翻译
- 九成生物医学论文有 AI 辅助写作痕迹_全翻译