Measuring AI-Generated Writing on arXiv: Trends and Limitations
Measuring AI-Generated Writing on arXiv: Trends and Limitations
對 12,750 篇 arXiv 論文的研究顯示,目前約有三分之一的新投稿被判定為機器寫作,且在 ChatGPT 發布後出現急劇增長。研究結果顯示學術寫作模式發生了重大轉變,儘管作者強調,檢測是識別「機器風格」的散文,而非證明特定的作者身份。
AI Prevalence Trends in Academic Papers
arXiv 論文中的機器寫作訊號在 2021 年和 2022 年期間維持在 0.4% 的平穩狀態,隨後分兩波上升,到最近一個完整季度約達到 32%,並在 2026 年初達到近 39% 的峰值。為了確保準確性,研究人員使用 pre-LLM 論文作為真實的人類控制組來校準檢測器,將閾值設定在僅有 0.4% 的真實 pre-ChatGPT 文本會被標記的情況下。
Distribution Across Scientific Fields
AI 風格寫作的採用程度因學科而異。在截至 2026 年 7 月之前的 12 個月內,被標記的比例分別如下:
| Field group | Pre-LLM control | Recent flagged share | 95% CI |
|---|---|---|---|
| Computer science | 0.2% | 65.0% | [59.3, 70.3] |
| Quantitative biology | 3.5% | 56.3% | [51.0, 61.7] |
| Electrical eng. & systems | 1.7% | 51.3% | [46.0, 57.0] |
| Economics & finance | 2.5% | 47.0% | [41.3, 52.7] |
| Applied physics | 1.3% | 34.0% | [29.0, 39.7] |
| Statistics | 1.8% | 31.3% | [26.0, 36.7] |
| Condensed matter | 0.0% | 24.0% | [19.3, 29.0] |
| High-energy physics | 0.5% | 14.0% | [10.0, 18.0] |
| Astrophysics | 0.0% | 10.7% | [7.3, 14.3] |
| Mathematics | 0.0% | 0.7% | [0.0, 1.7] |
Computer science 顯示出最高的盛行率,為 65%,而 mathematics 顯示出最低的,為 0.7%。
Methodological Framework and Limitations
該研究分析了 12,750 篇論文的 version-1 PDF,以防止修訂版本將現代文本洩漏到較舊的時間段中。研究人員對全文進行評分,而非僅針對摘要,因為摘要往往會低估 AI 訊號。
Key Limitations
- Mathematics Blind Spot: mathematics 的低分可能反映了檢測器的盲點,而非低採用率。由於數學論文主要由符號和定理-證明結構組成,其剩餘的散文部分較為稀疏,且與用於訓練檢測器的科學英語不同。
- Lower Bound Estimation: 由於檢測器可能無法涵蓋作者使用的每種私人的模型與提示詞組合,報告的盛行率被視為 AI 輔助寫作的真實比例的下限。
- Detection vs. Authorship: 標記表示文本「讀起來」像機器寫作;它無法區分完全生成的文檔,還是經過重度 AI 輔助編輯的文檔。
Community Insights and Counterpoints
研究人員與開發者的討論突顯了關於學術界使用 AI 以及檢測工具可靠性的幾種細微差別。
The Role of AI in Non-Native English Writing
許多研究人員建議,AI 主要被用作翻譯和潤色工具,而非內容生成器。
"I think most of the papers we write would be flagged by AI detectors... because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing."
Skepticism of Detection Accuracy
一些用戶回報了他們在 LLM 存在之前多年手寫的論文,卻得到了高 AI 評分,這表明儘管研究進行了校準,仍存在誤報的可能性。
"I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine."
其他批評者認為,學術寫作的性質——通常冗長且包含大量模板化內容——自然地模仿了 LLM 訓練時產生的模式,使得僅憑語言學分析,從根本上很難區分有機文本與合成文本。
Potential for Increased Research Output
一些人認為,LLMs 可能會降低研究人員的進入門檻,對於那些擁有強大技術能力但卻在學術寫作上感到困難的研究人員來說,這可能增加發表高品質研究的數量。