Relative Positional Encodings in Transformers: How Shaw's Attention, Transformer-XL, and T5 Relative Biases Preserve Translation Invariance
Relative Positional Encodings in Transformers: How Shaw's Attention, Transformer-XL, and T5 Relative Biases Preserve Translation Invariance Standard self-attention operations in transformer architectures possess no inherent awareness of sequence order. Because scaled dot-product attention computes interactions across sets of tokens without regard to index ordering, early models relied on absolute positional encodings to inject sequential structure. While absolute encodings assign rigid coordina
1 min
