AI 生成的劣质内容已无处不在,甚至渗透进了同行评审的学术文献。虚假引用、未经编辑的提示词回复,以及毫无意义的图表,都成功躲过了编辑和同行评审员的审查,而且目前尚不清楚相关责任人是否会面临任何后果。
如今,多个科学领域似乎将在同行评审或期刊介入之前,就对 AI 引发的问题执行规定。物理与天文学预印本服务器 arXiv 的一位相关人士通过社交媒体发帖宣布,任何向该服务器提交的不当 AI 生成内容,将导致提交者被禁言一年,并且未来所有投稿在 arXiv 托管之前,必须永久性地先经过同行评审。
托马斯·迪特里希除了是俄勒冈州立大学的荣休教授外,还深度参与 arXiv 的工作,担任其编辑咨询委员会成员和审核团队成员。因此,他完全有能力理解该组织的政策,不过我们也已联系 arXiv 领导层寻求确认,但尚未收到回复。
在 X 平台(对于没有 X 账号的用户,相关内容也在 Bluesky 上被截图)的一个帖子中,迪特里希将这项新政策描述为直接源于 arXiv 的审核标准。这些标准写道:“提交至 arXiv 的内容必须在形式上符合学术交流的适当标准,包括恰当且精心准备的章节、图表、表格、参考文献等。要求稿件准备过程普遍严谨细致。”
迪特里希还指出,手稿的所有作者均对其内容负责。因此,如果他们粗心大意地提交了由 AI 生成且违反这些准则的材料——迪特里希列举了“不当语言、剽窃内容、偏见内容、错误、谬误、不正确的参考文献或误导性内容”——那么责任在于作者,而非 AI。一旦发现违规行为,手稿列出的所有作者将面临一年的提交禁令,并且未来的任何手稿只有在经过期刊同行评审后,才会被 arXiv 接受。
对于严重依赖 arXiv 的领域来说,这些是严厉的制裁。在天体物理学等领域发布预印本被广泛视为正常发表流程的一部分,科学家们通常会从预印本获得反馈,从而改进他们提交同行评审的内容。不幸的是,与其他大多数事物一样,这个系统也可能被钻空子——人们可以提交有缺陷的内容,并将从未参与过的人列为作者。幸运的是,其审核系统包含一个申诉程序。
当在出版物中发现这些问题时,一个显而易见的问题是为什么没有人更早发现它们。现在,我们至少知道有人在尝试这样做。
周一,arXiv 管理层向 Ars 证实,这是该组织的官方政策。
AI-generated slop has shown up everywhere, including in the peer-reviewed literature. Fake citations, unedited prompt responses, and nonsensical diagrams have all slipped past editors and peer reviewers, and it’s not always clear if there are any consequences for the people responsible.
Now, it appears that a number of scientific fields will be enforcing rules against AI-generated problems even before peer review or journals get involved. One of the people involved in the physics and astronomy preprint server arXiv used a social media thread to announce that any inappropriate AI-produced content submitted to the server will result in a one-year ban and a permanent requirement that future publications undergo peer review before the arXiv will host them.
Thomas Dietterich, in addition to being an emeritus professor at Oregon State University, is heavily involved with arXiv, serving on its editorial advisory council and on its moderation team. So he’s in a good position to understand the organization’s policies, although we have also reached out to arXiv leadership for confirmation, but have not yet received a response.
In a thread on X (also screenshotted on Bluesky, for those without X accounts), Dietterich described the new policy as arising directly from the arXiv’s moderation standards. “Submissions to arXiv must comply with appropriate standards of scholarly communication in form, including appropriate and carefully prepared sections, figures, tables, references, etc.,” those standards read. “General scrupulousness and care of preparation are required.”
Dietterich also notes that all authors of a manuscript are responsible for its content. So, if they carelessly submit material generated by an AI that violates these guidelines—Dietterich cites “inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content”—then they’re responsible, not the AI. Should violations be discovered, all of the manuscript’s listed authors will now receive a one-year submission ban, and any future manuscripts will only be accepted after they’ve been through peer review by a journal.
For fields that rely heavily on the arXiv, those are severe sanctions. Posting preprints in areas like astrophysics is widely considered part of the normal publication process, and scientists will often get feedback on preprints that helps them improve what they submit for peer review. The unfortunate problem is that, like most other things, the system can be gamed—people could submit flawed content that lists people as authors who have never been involved. Fortunately, its moderation system includes an appeal process.
One obvious question that arises when these problems are found in publications is why nobody caught them sooner. Now, we can at least know that someone is trying to.
On Monday, arXiv leadership confirmed to Ars that this is the organization’s official policy.