跳到正文
Nathan Lambert· @natolambert · X·· 3 小时前AI 评分30
AI 导读

Nathan Lambert 表示正为 Mercor 提供研究方向的建议,主张构建可扩展的样本高效 RL 数据方法,并推动评估科学前沿、针对明确经济价值领域建立专门 benchmark。他认为前沿实验室将围绕真实世界代表性评估与合成数据方法大举投入,并看好基于开源模型构建这一方向的机会。

正文

I'd frame it as follows: The data industry has taken off in recent years, but the quality of our net output is still far too low (amplifying behaviors like reward hacking).

We're working to build scalable methods for creating sample-efficient data for RL. In order to keep this pipeline going, we need to push the frontier of evaluation science, while building specific benchmarks to hillclimb on areas of clear economic value.

I've been advising Mercor on how to build this research direction effectively. These are my views, but I'm confident we're going to see a major investment from economy around the frontier labs (open inference, open post-training, and data) orient around expertise in building real-world representative evals and synthetic data methods to scaling training data around them.

I’m personally very excited about this, it is the research that will make more of the economy “feel the AGI” for the first time.
(And, there’s a big opportunity to build this on open models.)

引用Edward Hu@edwardjhu
Mercor is building a world-class research team. As the leading AI data provider, we are uniquely positioned to combine benchmarks, data production, model training, and economics research to advance model productivity. We are committed to sharing our findings with the world. DM me if this sounds exciting. https://www.mercor.com/blog/why-mercor-is-building-a-research-team/
在 X 查看被引用的帖子

来源:Nathan Lambert · x.com