SemiAnalysis· @SemiAnalysis_ · X·· 2 小时前AI 评分41
AI 导读
我们在 ExploitGym 上评估了 GLM-5.3 的网络能力,并分析了执行轨迹。GLM-5.3 将大部分执行预算用于测试隐藏的运行时条件是否改变了其结论。在问题 arvo5665 中,GLM-5.3 通过 sanitizer 构建、语料测试和目标 fuzzing 探索了更多周边程序。(1/3)🧵 https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
正文
We assessed GLM-5.3 cyber capabilities on ExploitGym and analyzed the traces. GLM-5.3 spent much of its execution budget testing whether hidden runtime conditions changed its conclusion. In problem arvo5665, GLM-5.3 explored more of the surrounding program through sanitizer builds, corpus tests, and target fuzzing. (1/3)🧵
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
来源:SemiAnalysis · x.com