Is the @METR_Evals long horizons measure effectively saturated?
No updates to the most famous graph of AI progress since May, but other published work suggests you can effectively get 18+ weeks of work out of Fable (and I assume GPT-6) in harnesses.
@METR_Evals 的长时程指标是否实际上已经饱和? 自5月以来,AI进展最著名的图表没有更新,但其他已发表的研究表明,在受控环境中,Fable(我推测GPT-6也是如此)可以有效完成18周以上的工作量。
@METR_Evals 的长时程指标是否实际上已经饱和? 自5月以来,AI进展最著名的图表没有更新,但其他已发表的研究表明,在受控环境中,Fable(我推测GPT-6也是如此)可以有效完成18周以上的工作量。
Is the @METR_Evals long horizons measure effectively saturated?
No updates to the most famous graph of AI progress since May, but other published work suggests you can effectively get 18+ weeks of work out of Fable (and I assume GPT-6) in harnesses.
来源:Ethan Mollick· x.com