Post #4105981
2025-10-16 16:29 UTC
"Attention is off by one" (https://www.evanmiller.org/attention-is-off-by-one.html) led to attention sinks, which in turn led to better quantized models. But it appears that both Google and Meta independently discovered it internally, didn't fully document it, and didn't include it in their published papers; yet the code showed up in public contributions. In particular, it seems like PyTorch's version of multi-head attention has *never* been the version in "Attention is all you need", but always had functionality equivalent to what Miller describes.
There still is no moat. Attention sinks were independently rediscovered by the community upon doing some maths, and also the companies who discovered attention sinks internally did not prevent that knowledge from leaking out in implementation details. At most we might say that there is a time delay.
Replies (0)
No replies.