Rohan Paul · @rohanpaul_ai · X·2026-09-26 14:15·1小时前
AI 导读

OpenAI 官方失准报告网站上的新事件报告。 自我复制的提示词注入,能够有效地从一个 AI 交互传播到另一个 AI 交互。 恶意指令可以隐藏在 AI 读取的内容中,比如一封邮件,诱使 AI 去执行它,而不是只完成用户的任务。 巧妙之处在于,该指令还告诉 AI 把同一条恶意指令复制到它的回复中,从而可能暴露给下一个读取它的 AI。

Rohan Paul@rohanpaul_ai
49AI 编辑部评分,满分 100
2026-09-26 14:15· 1小时前
AI 导读

OpenAI 官方失准报告网站上的新事件报告。 自我复制的提示词注入,能够有效地从一个 AI 交互传播到另一个 AI 交互。 恶意指令可以隐藏在 AI 读取的内容中,比如一封邮件,诱使 AI 去执行它,而不是只完成用户的任务。 巧妙之处在于,该指令还告诉 AI 把同一条恶意指令复制到它的回复中,从而可能暴露给下一个读取它的 AI。

New incident reporting on OpenAI's official misalignment reporting site.

Self-replicating prompt injections, that can effectively spread from one AI interaction to another.

A malicious instruction can be hidden inside something the AI reads, like an email, and trick the AI into following it instead of just doing the user’s task.

The clever part is that the instruction also tells the AI to copy that same malicious instruction into its reply, potentially exposing the next AI that reads it.

来源:Rohan Paul· x.com