Chelsea Finn· @chelseabfinn · X·· 2 小时前AI 评分34
AI 导读
随着机器人变得越来越可靠,DAgger 的效率就大大降低了。 例如 90% 的成功率意味着人 >90% 的时间都在被动地看着机器人 核心思路:失败来自相似的初始状态,我们可以把学习聚焦在这些状态上! 论文:https://arxiv.org/abs/2610.05882
正文
As robots become more reliable, DAgger becomes far less efficient.
eg 90% success means the person spends >90% of the time passively watching the robot
Key idea: failures come from similar initial states & we can focus learning on these states!
I've been thinking a lot about making robot learning scalable. To me, that means learning on the job, and using the human help still needed efficiently. In Mulligan ⛳️, we focus each data-collection round on the initial states where the robot fails. This happens in deployment, with the robot doing the task and a person stepping in when needed. On 3 real tasks and 2,550 blind evals, it improves on uniform collection at the same budget.在 X 查看被引用的帖子
来源:Chelsea Finn · x.com