Base Labs 研究:单步梯度更新的"不合理效率"——RL 训练轨迹的线性结构可被利用
The Unreasonable Efficiency of a Single Gradient Step
Base Labs 复现了 RL 优化轨迹近似"直线"的结论,发现将 Qwen2.5-1.5B 在 MATH 上的首个梯度步放大 1000 倍,效果超过 500 步 RL 训练的模型(66.7% vs 63.8%)。
Rollout · OCT 2026
Our in progress work to better understand how to exploit linear structure in RL trajectories

Overview
RL has become a necessary component for training frontier language models, but is also is incredibly expensive. If the data, problems, and environments LLMs learn over are highly structured, then surely the weight dynamics for RL training are structured in some exploitable way too. A number of recent papers seem to indicate a rather striking structure result on RL optimization: they are effectively straight. Extrapolating from an early stretch of RL training can recover much of the improvement from a substantially longer run [1, 2, 3].
While many of these papers differ in their specific construction of a straight or low-rank trajectory, they paint a unifying picture: the same gains from a full RL run can be achieved by linearly extrapolating a few early optimization steps. This seems almost too good to be true.
In this short release update we replicate this claim and find that unfortunately, in some crucial ways, it kind of is. We find:
Scaling even just the first gradient step 1000x with Qwen2.5-1.5B on MATH exceeds a 500-step RL’d model (66.7% vs 63.8%).
The observed improvements largely comes from fixing one failure mode, namely fixing where the base model doom loops.
This effect weakens when off-distribution. For example, first step scaling does not help function calling or Knights & Knaves.
While our work requires further development, we believe that exploiting such linear structure in RL trajectories may (in certain contexts) be about removing easy failure modes rather than gaining fundamentally new capabilities. After all, there is rarely a ‘free lunch’.
Cite this work
@article{baselabs2026-the-unreasonable-efficiency-of-a-single-gradient-step,
title = {The Unreasonable Efficiency of a Single Gradient Step},
author = {Psenka, M. and Shariff, A. and Kirkby, M.},
journal = {Base Labs},
year = {2026},
number = {001},
url = {https://labs.baseten.co/001},
}Related from the lab
来源:Baseten Base Labs:模型研究 · labs.baseten.co