Tibo · @thsottiaux · X·2026-09-09 08:28·42分钟前
AI 导读

Anthropic 高管称 Astra 是其迄今最强大且对齐最好的模型,并指约两个月前已在实际中解决 Claude 模型的提示注入问题。其引用推文显示 OpenAI 新模型在提示注入风险上与 Gemini Flash 和 Opus 4.8 大致相当。他呼吁行业投入更多精力训练模型抵抗提示注入,并称将持续评估其他实验室以推动安全对齐。

Tibo@thsottiaux
38AI 编辑部评分,满分 100
2026-09-09 08:28· 42分钟前
AI 导读

Anthropic 高管称 Astra 是其迄今最强大且对齐最好的模型,并指约两个月前已在实际中解决 Claude 模型的提示注入问题。其引用推文显示 OpenAI 新模型在提示注入风险上与 Gemini Flash 和 Opus 4.8 大致相当。他呼吁行业投入更多精力训练模型抵抗提示注入,并称将持续评估其他实验室以推动安全对齐。

I for one am glad prompt injection is getting solved across the industry and that Astra is not only our most capable, but also our most aligned model to date. But also ...

Boris ChernyI am pleased to see that OpenAI’s new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work! Evaluating and naming other la...

来源:Tibo· x.com