Nathan Lambert · @natolambert · X·2026-09-22 05:41·1小时前
AI 导读

强烈赞同。这种情况比我预期的要少得多——部分原因是构建高质量环境通常涉及相当昂贵的验证(测试强模型)。尽管如此,还可以有更多。 [@Thom_Wolf]:发布大量高质量的开源RL环境是当前任何人推动开源前沿所能做的最有影响力的事情 相当于分享高质量预训练数据,但在新的RLVR范式下

Nathan Lambert@natolambert
27AI 编辑部评分,满分 100
2026-09-22 05:41· 1小时前
AI 导读

强烈赞同。这种情况比我预期的要少得多——部分原因是构建高质量环境通常涉及相当昂贵的验证(测试强模型)。尽管如此,还可以有更多。 [@Thom_Wolf]:发布大量高质量的开源RL环境是当前任何人推动开源前沿所能做的最有影响力的事情 相当于分享高质量预训练数据,但在新的RLVR范式下

Strong agree. There’s way less of this than I would expect — partially due to the fact that making a high quality environment usually involves fairly expensive verification (testing strong models). Still, there can be much more.

Thomas Wolfreleasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of s...

来源:Nathan Lambert· x.com