Post #3077905
2025-06-09 22:17 UTC
@glyph @xgranade I do feel like we have already progressed quite a bit from the original transformers paper(rotary embeddings, sparse/chunked attention,MoE(which I frankly don't understand myself)) seems to be standard in every new model.
Then there's the whole field of "linear attention" (RWKV my beloved ♥️), which is constantly catching up and has a lot of research directions and hacks similar to pre-transformer era deep learning architectures
Replies (0)
No replies.