跳到正文
Zyphra Research·· 19 小时前AI 评分51

Zyphra 解析混合 NoPE 模型如何在没有位置编码的情况下学到位置

How Hybrid NoPE Models Learn Position Without Position Encodings

AI 导读

Zyphra 发文解释混合架构中全局 NoPE 注意力如何获得位置信息:滑动窗口注意力等局部混合层在残差流中留下随相对距离变化的近因偏置结构,训练中全局层的 query/key 投影学会读取该结构,形成隐式相对位置编码。

正文
Introduction

Treatment of position is a fundamental consideration in the design of Transformer language models. Attention, considered on its own, does not care about the order of its inputs. That is, without some way to distinguish positions, swapping two tokens would not change the basic attention operation. Natural language obviously does care about order: “the dog chased the cat” and “the cat chased the dog” contain the same tokens but mean something very different.

For this reason, Transformers typically add an explicit position encoding that tells attention where tokens occur in a sequence. Today, one of the most common approaches is rotary position encoding, or RoPE, which modifies queries and keys according to their positions before computing attention.

Explicit position encodings work extremely well, but they come with tradeoffs. In particular, positional mechanisms learned over one context length can behave poorly when extrapolated far beyond it. RoPE-based models, for example, can require special long-context training or interpolation techniques when extending their context window. NoPE models remove explicit position encodings entirely and can thus behave better when extrapolating in length. However, historically, they have generally underperformed models with explicit positional information within the training distribution since they cannot properly represent the vital positional information in language. 

Recently, however, an interesting middle ground has emerged. Modern hybrid architectures increasingly interleave local mixing layers—such as sliding window attention or gated linear attention—with occasional global attention layers. And, increasingly, these global layers use NoPE.

That creates a puzzle: if global NoPE attention receives no explicit positional information, how does it know where tokens are? Our work provides an answer.

The Missing Position Encoding Is Already in the Residual Stream

The core idea is surprisingly simple. A local layer mixes information differently depending on how far apart two tokens are. As a result, it leaves a signature of relative position directly in the geometry of the model’s hidden representations. 

Consider sliding window attention. Instead of allowing each token to attend to the entire preceding sequence, SWA restricts attention to a fixed-size local window. Imagine, for intuition, that a sliding window simply averaged the representations within that window. Two adjacent tokens would then be averages over almost exactly the same set of preceding tokens. Move them slightly farther apart and their windows overlap less. Move them more than one window apart and they may not overlap at all.

So the closer two positions are, the more information their representations share.The result is a recency bias in the residual stream: nearby token representations are, on average, more similar than distant ones.

Figure 1: Tokens initially have no positional relationship. A local mixing layer correlates nearby residual states. Global NoPE attention then learns query and key projections that convert this lag-dependent residual structure into recency-biased attention logits.

This is important because a recency bias alone provides information about relative position. The model does not necessarily need a variable explicitly saying “this key is 37 tokens behind the query”. It can instead learn that progressively less similar residual states generally correspond to progressively greater relative distances.

Many conventional relative position encodings ultimately create a similar effect in attention: they make recent keys systematically easier to attend to than distant ones. Our result shows that hybrid models can create this structure implicitly, without explicitly injecting relative positions into global attention.

Step 1: Local Mixing Creates a Recency Bias

The first part of our mechanism is structural. Since sliding window attention behaves roughly like a content-dependent local filter, nearby outputs are constructed from many of the same inputs, so they become correlated. The amount of overlap falls as relative distance increases.

Critically, the scale of this effect is determined by the window size, not the total sequence length. This distinguishes SWA from the weak implicit positional signal that can arise in a Transformer using global NoPE attention everywhere.

With global causal attention, a token late in a sequence aggregates information from almost everything before it. Two nearby positions late in a long sequence therefore have almost identical histories. For example, around position 1,000, representations ten positions apart can still share roughly 99% of the tokens included. As the sequence grows, fixed relative distances become increasingly difficult to distinguish.

With sliding window attention, the comparison is very different. A 128-token window remains a 128-token window whether the overall sequence contains 1,000 or 100,000 tokens. The local overlap pattern therefore does not get washed out simply because the sequence becomes longer.

We see exactly this behavior experimentally. At initialization, global NoPE attention produces representations whose fixed-distance differences become difficult to resolve even for moderate sequence lengths. SWA instead maintains a strong distance-dependent similarity profile throughout its local window, independent of where the query appears in the sequence.

In other words, locality turns out to provide more than computational efficiency. It also creates a useful positional inductive bias.

The Signal Gets Stronger During Training

An architectural bias at initialization would not be particularly useful if training immediately erased it. We find the opposite.

As training progresses, the recency bias in the residual stream strengthens. It also generally becomes stronger deeper in the network, suggesting that this positional structure accumulates as information repeatedly passes through local mixing layers.

Figure 2: Similarity as a function of relative distance at early, middle, and late SWA layers, showing that the recency structure appears from initialization and strengthens during training and through depth.

This happens early in training and remains robust throughout. We observe the same qualitative behavior across both the 120M- and 350M-parameter models studied. This should not be surprising since having a mechanism to represent relative position information seems fundamental to good language modelling performance. 

We also vary the size of the sliding window from 64 to 4096 tokens. As the mechanism predicts, smaller windows generally create a sharper and stronger recency signal. This remains true even when we remove RoPE from the sliding window layers themselves, demonstrating that the phenomenon does not require an explicit position encoding inside SWA.

This gives us the first half of the story: local mixing writes relative position information into the residual stream. But that alone is not sufficient. Global attention still needs some way to use it.

Step 2: Global NoPE Attention Learns to Read It Out

A global NoPE layer has no direct position encoding. But, it turns out to be enough that its attention logits come from the interaction between its query and key projections of the residual stream.

Because of the preceding local layers, the residual stream now contains a systematic pattern in which representations depend on relative distance. The global attention layer can learn query and key projections that align with precisely this part of the residual geometry.

One way to think about it is that the local layer writes a positional signal in a high-dimensional representation, while global attention learns the directions along which to read that signal.

At random initialization, the query and key projections have no particular reason to align with the recency structure, so the global NoPE attention logits initially have relatively little distance dependence.

During training, however, this changes. The query and key matrices become aligned in a way that turns the recency structure in the residual stream into a recency bias in the attention logits. Recent keys receive systematically different scores from distant ones, despite there being no explicit positional term in the global attention operation.

And that is exactly what we observe.

Figure 3: At initialization, attention logits have little systematic dependence on distance. During training, a clear recency-biased profile emerges as the global heads learn to select the positional structure already present in the residual stream.

The effect emerges early in training and becomes stable across global NoPE layers. Rather than independently inventing a position encoding from the causal mask, these layers appear to make use of positional structure created upstream by the model’s local mixers.

So the complete mechanism is:

Local mixing → distance-dependent residual geometry → learned query/key alignment → implicit relative position encoding in global NoPE attention.

Why This Is Different From Ordinary NoPE

Previous work has shown that a model built entirely from global NoPE attention can recover some notion of position from the causal mask. But there is an important limitation.

The causal mask tells every token which tokens precede it. In early positions, that asymmetry can produce a useful position-dependent signal. Far into a sequence, however, adjacent tokens both have access to almost the same enormous prefix. The difference between their histories becomes proportionally tiny. The resulting positional signal therefore loses resolution as sequence length increases.

The hybrid mechanism we identify behaves differently. Its scale is anchored by the local mixer. For SWA, that scale is the window size. For a recurrent gated linear layer, it can instead be set by the rate at which historical information decays. In either case, it is a local scale that does not grow with total sequence length.

That provides a natural explanation for why hybrid architectures can combine local layers with global NoPE attention without leaving the global layers position-blind. The global layers inherit a relative positional coordinate system whose resolution does not inherently disappear as context length grows.

Window Size Provides a Useful Test of the Mechanism

If local mixing is responsible for creating the positional signal, changing how local the layer is should change the strength of that signal. It does.

Across our SWA models, smaller windows create substantially stronger residual-stream recency biases. In models where SWA itself also uses NoPE, this effect is especially clear.

Interestingly, we also find a corresponding trend in language-modeling performance: smaller-window models achieve lower validation loss than their larger-window counterparts in these experiments. For NoPE-in-SWA models, sufficiently large windows can lead to training collapse.

This result should be interpreted carefully. Changing the window size changes more than just the strength of recency bias, and our experiments are not designed to isolate recency as the sole causal factor behind the loss differences.

But the pattern is consistent with the mechanism. If the local layer becomes too global, its ability to create a well-resolved positional structure weakens. The subsequent global NoPE layers then have less useful relative-position information available to read out.

Figure 4: Smaller windows produce stronger recency structure in the residual stream and lead to lower language modeling loss, with or without RoPE within the SWA window. In the NoPE-in-SWA setting particularly, excessively large windows also coincide with instability in the learned global-logit recency signal and a collapse in language-modeling performance.

The Mechanism Is Not Specific to Sliding Window Attention

If the important ingredient is locality or decaying memory, rather than SWA itself, the same phenomenon should appear with other sequence mixers. We test this using KDA, a gated linear attention mechanism proposed by Kimi in Kimi-Linear.

Unlike SWA, KDA does not use a hard fixed window. Instead, information from previous positions is repeatedly mixed through a recurrent state and gradually decays. But the same intuition still applies – nearby outputs share more historical information than distant outputs and thus become correlated. In a simplified mathematical model, this gives an exponentially decaying similarity profile rather than the roughly triangular profile produced by uniform sliding window attention.

Empirically, we observe the same overall mechanism in KDA–NoPE hybrids: KDA layers increase residual-stream recency structure through training and depth, and subsequent global NoPE layers develop recency-biased attention logits. This suggests that the phenomenon is more general than a particular attention implementation. Local and recurrent sequence mixers can create positional structure that later global attention layers exploit.

Why This Matters

Hybrid architectures are increasingly attractive for long-context language models. Local and linear layers reduce the cost of processing long sequences, while occasional global attention layers preserve the ability to connect information across the entire context. Our results show that this division of labor may have another benefit: the local layers can provide the positional coordinate system that the global layers need.

This gives a mechanistic explanation for how global NoPE attention can work inside architectures where, considered in isolation, it appears to be missing something fundamental.

It also suggests an interesting direction for long-context modeling.

Explicit position encodings such as RoPE define a positional transformation that can behave very differently outside the range encountered during training. By contrast, the positional mechanism we identify is based on local residual-stream geometry. A fixed local window does not change merely because the model is processing a much longer sequence, and a recurrent decay process similarly has a characteristic scale independent of total context length.

The model can therefore maintain a well-resolved local notion of relative position while its global NoPE layers remain free to attend across arbitrary distances.

We have not yet established a quantitative relationship between this mechanism and downstream long-context extrapolation performance, and understanding that connection is an important direction for future work. But the mechanism gives us a concrete hypothesis for why hybrid NoPE architectures are particularly promising for context extension – their position encoding is implicit, relative, and anchored to a length-independent local operation rather than to absolute sequence length.

More broadly, this changes how we can think about position in Transformers. Positional information does not necessarily need to be attached explicitly to every attention operation. It can instead emerge as a property of the model’s internal representation—written into the residual stream by one class of layers and selectively read by another.

That opens up a larger design space for long-context architectures: rather than asking only which position encoding should we add to attention, we can ask what geometry should the architecture create so that attention can infer position for itself?

来源:Zyphra Research · zyphra.com