跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分36
AI 导读

Meta 首席 AI 官 Alexandr Wang 表示,AI 对齐目前仍是科学上最开放的难题之一,没人知道确切解法。他提出的思路是"可扩展监督":用不同的 AI 去观察并约束更聪明的 AI,监督方也需随之不断升级,即打造"越来越聪明的监管智能体"。Meta 的 Muse 已采用该方案的一个版本,用独立的哨兵智能体检查主智能体的行为。

正文

Alexandr Wang ( @alexandr_wang, Chief AI Officer at Meta): nobody knows how to solve alignment yet

"This is, I think, one of the most open questions scientifically in AI. I think nobody knows exactly the way to solve this problem, but there’s a few ideas."

His solution is "scalable oversight":

“as the AIs get smarter, we use a different set of AIs to observe what they’re doing and keep them in check.”

The watcher AIs have to improve along with the models they watch, so labs would need to build “smarter and smarter policing agents” too.

Meta’s Muse already uses a version of this, with a separate sentinel agent checking what the main agent does.

----
Full video on "Cleo Abram" YouTube channel, (link in comment)

来源:Rohan Paul · x.com