跳到正文
Arena.ai· @arena · X·· 1 小时前AI 评分44
AI 导读

今天在 @tbpn 上收听 @ml_angelopoulos 带来的另一期 Arena Alignment Index 解读

正文

Hear another Arena Alignment Index breakdown from @ml_angelopoulos on @tbpn today

引用TBPN@tbpn
Arena's @ml_angelopoulos says one of the hardest problems in alignment is that users often can't tell when the model is helping them or hurting them. "You really need alignment signals that are independent. We have three at Arena." "The first is that models take unauthorized actions, which means that you permission them in a certain way, but they break those permissions." "It can cause things like the Hugging Face incident. But it can also cause very mundane problems like losing important information within your company or on your own laptop." "The second signal is called deceptive completion." "The model will tell you that it did something, but it didn't actually do it." "Third is false attribution. It'll attribute intent to the user when that intent was not supposed to be there." "If models are able to do this, then certainly they're not perfectly safe for a user and they're not perfectly aligned to user intent. They might actually hurt people down the line."
在 X 查看被引用的帖子

来源:Arena.ai · x.com