Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
SAE 潜在空间中的词性作为涌现类别
AI 导读
研究以词性(PoS)为受控测试用例,考察 Sparse AutoEncoders(SAE)的潜在表示暴露何种语言结构。结果显示词性区分可从 SAE 激活中高度恢复,但不与单个 latent 一一对应,且不能还原为词汇记忆;开放与封闭词类差异显著,各类别由紧凑的稀疏 latent 组支撑,这些组在留出数据上保持稳定,相关类别间存在重叠。
HuggingFace Daily Papers(社区热门论文)
36
AI 编辑部评分,满分 100SAE 潜在空间中的词性作为涌现类别
研究以词性(PoS)为受控测试用例,考察 Sparse AutoEncoders(SAE)的潜在表示暴露何种语言结构。结果显示词性区分可从 SAE 激活中高度恢复,但不与单个 latent 一一对应,且不能还原为词汇记忆;开放与封闭词类差异显著,各类别由紧凑的稀疏 latent 组支撑,这些组在留出数据上保持稳定,相关类别间存在重叠。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org