Transformer 在残差流中提前编码对话伙伴专业水平,但到网络中点后才真正影响输出

HuggingFace Daily Papers(社区热门论文)·2026-09-07 08:00·2天前
AI 导读

一项新研究发现,Transformer 能在残差流中提前线性解码对话伙伴的专业水平属性,但该属性要到网络中点之后才真正影响输出。基于多轮研究规划对话语料库 ExpertCollab 的实验显示,伙伴专业水平在早期层最易解码,中点后降至接近随机水平;反事实修补表明,在峰值可解码层注入差异几乎不影响后期读出,而中点后注入则几乎完全传播,两者差距超一个数量级。

HuggingFace Daily Papers(社区热门论文)
37AI 编辑部评分,满分 100

Transformer 在残差流中提前编码对话伙伴专业水平,但到网络中点后才真正影响输出

2026-09-07 08:00· 2天前
AI 导读

一项新研究发现,Transformer 能在残差流中提前线性解码对话伙伴的专业水平属性,但该属性要到网络中点之后才真正影响输出。基于多轮研究规划对话语料库 ExpertCollab 的实验显示,伙伴专业水平在早期层最易解码,中点后降至接近随机水平;反事实修补表明,在峰值可解码层注入差异几乎不影响后期读出,而中点后注入则几乎完全传播,两者差距超一个数量级。

Abstract:A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network. Counterfactual patching shows that injecting the expertise difference at the layer of peak decodability barely changes a fixed late-layer readout, whereas the same difference injected past the midpoint propagates almost completely, a separation of more than an order of magnitude. A content-matched random control and a probe-free diagnostic place the transition at the same early layer, and a statically specified control attribute stays decodable throughout. An inferred relational attribute is therefore represented well before it becomes causally active, which bounds where any attempt to read out or steer partner-conditioned behavior must intervene. We use one model on a synthetic corpus as an initial demonstration.
Comments:
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.07139 [cs.AI]
  (or arXiv:2609.07139v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.07139
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mika Okamoto [

Mon, 7 Sep 2026 07:41:57 UTC (173 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org