Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分37
AI 导读
你可以预测 LLM 内部状态的走向,并且有针对性的编辑可以将其拉回正轨。 在一个冻结的 Hermes 8B 上,从模型某一层的隐藏状态中取出的 8 个数字,可以预测几层之后发生的 68 项测量结果。 预测误差比简单的平均基线低约 69-76%。
正文
You can predict where an LLM's internal state is heading, and that targeted edits can pull it back on course.
On a frozen Hermes 8B, 8 numbers, taken from the model's hidden states at one layer, predict 68 measurements of what happens several layers later.
The prediction error comes out about 69-76% lower than a simple average baseline.
We show that a small set of internal coordinates can predict future hidden states—and, more importantly, that intervening on those coordinates measurably changes downstream internal dynamics. In our experiments, 8 internal coordinates predict a 68-dimensional future target, while targeted interventions reduce downstream internal prediction error by 21.68–43.79% across registered cells. The broader idea is simple: If you can predict where a model's internal state is going, can you intervene and change where it goes? These results suggest that the answer may be yes. https://proprioceptiveai.com/From_Internal_Prediction_to_Measured_Control.pdf在 X 查看被引用的帖子
来源:Rohan Paul · x.com