Measuring AI-Generated Writing on arXiv: Trends and Limitations

Measuring AI-Generated Writing on arXiv: Trends and Limitations

对 12,750 篇 arXiv 论文的研究表明,目前大约三分之一的新提交论文读起来像是机器编写的,且在 ChatGPT 发布后出现了剧增。研究结果表明学术写作模式发生了重大转变,尽管作者强调,检测手段识别的是类机器风格的散文,而非证明特定的作者身份。

AI Prevalence Trends in Academic Papers

arXiv 论文中的机器编写信号在 2021 年和 2022 年期间一直稳定在 0.4%,随后经历了两次浪潮,到最近一个完整季度上升至约 32%,并在 2026 年初达到近 39% 的峰值。为了确保准确性,研究人员使用 pre-LLM 论文作为人类对照的基准真相(ground-truth),设定了一个阈值,使得只有 0.4% 的真实 pre-ChatGPT 文本会被标记为 AI 编写。

Distribution Across Scientific Fields

AI 风格写作的采用情况因学科而异。在截至 2026 年 7 月的 12 个月内,被标记的比例分别为:

Field group Pre-LLM control Recent flagged share 95% CI
Computer science 0.2% 65.0% [59.3, 70.3]
Quantitative biology 3.5% 56.3% [51.0, 61.7]
Electrical eng. & systems 1.7% 51.3% [46.0, 57.0]
Economics & finance 2.5% 47.0% [41.3, 52.7]
Applied physics 1.3% 34.0% [29.0, 39.7]
Statistics 1.8% 31.3% [26.0, 36.7]
Condensed matter 0.0% 24.0% [19.3, 29.0]
High-energy physics 0.5% 14.0% [10.0, 18.0]
Astrophysics 0.0% 10.7% [7.3, 14.3]
Mathematics 0.0% 0.7% [0.0, 1.7]

Computer science 显示出最高的流行度,为 65%,而 mathematics 显示出最低的,为 0.7%。

Methodological Framework and Limitations

该研究分析了 12,750 篇论文的 version-1 PDF,以防止修订版本将现代文本泄露到较早的时间段中。研究人员对全文进行评分,而不是摘要,因为摘要往往会低估 AI 信号。

Key Limitations

  • Mathematics Blind Spot: 数学领域的低分可能反映了检测器的盲点,而非采用率低。由于数学论文主要由符号和定理-证明结构组成,剩余的散文部分非常稀疏,且与用于训练检测器的科学英语不同。
  • Lower Bound Estimation: 由于检测器可能无法覆盖作者使用的每种私有的模型和提示词组合,报告的流行度被视为 AI 辅助写作的真实比例的下限。
  • Detection vs. Authorship: 标记意味着文本 读起来 像机器编写的;它无法区分完全生成的文档与经过重度 AI 辅助编辑的文档。

Community Insights and Counterpoints

研究人员和开发人员之间的讨论突出了关于在学术界使用 AI 以及检测工具可靠性的几个细微差别。

The Role of AI in Non-Native English Writing

许多研究人员建议,AI 主要被用作翻译和润色工具,而不是内容生成器。

"I think most of the papers we write would be flagged by AI detectors... because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing."

Skepticism of Detection Accuracy

一些用户报告称,他们在 LLM 存在之前多年手动编写的论文也获得了高 AI 评分,这表明尽管研究进行了校准,但仍存在误报的可能性。

"I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine."

其他批评者认为,学术写作的性质——通常是冗长且模板化的——自然地模仿了 LLM 训练时产生的模式,这使得仅通过语言学分析从有机文本中区分合成文本在根本上是困难的。

Potential for Increased Research Output

一些人认为,LLMs 可能降低了研究人员的技术门槛,对于那些拥有强大技术能力但难以应对学术写作的阻碍,这可能会增加发表的优质研究的数量。

Sources