#transformers

52 posts · Last used 8d

Back to Timeline
Bogdan Buduroiu @budududuroiu@hachyderm.io · Aug 02, 2026
Morning, today we're looking at #MechanisticInterpretability , a subfield of AI Safety that attempts to understand the inner workings of artificial intelligence by analysing concrete structures, algorithms and circuits. Why do we even need to do this? Because the hope of understanding neurons as being features died on polysemanticity -- models represent more features than dimensions by assigning them to an overcomplete set of non-orthogonal directions (i.e. you can't hope that concepts can be broken down into linear combinations of features). This isn't a new field, people have been projecting intermediate GPT layers through the final layer activation and looking at the top-k most likely next tokens since GPT-2. And that's how we arrive at the Natural Language Autoencoder -- an autoencoder over residual-stream activations where the bottleneck is natural language text instead of a sparse vector. Nothing in the training objective requires the verbalisation bottleneck to be readable, faithful, or even semantically related to the intermediate layer being investigated, so while the training algorithm is typical of autoencoders, two things are distinct here: 1) both parts of the autoencoder (here the activation-verbaliser, AV, and the activation-reconstructor, AR) undergo special SFT to ensure the verbaliser and reconstructor can generate and read natural language from intermediate layer activations 2) a KL-divergence penalty is baked into the objective to make sure that, during joint training, the NLA doesn't diverge from the SFT version A side-product of this training is that we can take the current verbalisation and "desired" verbalisation and compute steering vectors from their difference. Surprisingly, Anthropic didn't find any evidence that the NLA was engaging in steganography to smuggle information to avoid human detection, but it did find that NLAs tend to confabulate a lot. #AIResearch #Transformers #TransformerCircuits #Anthropic #NLA
1
0
0
devopscats @devopscats@toot.cat · Jul 28, 2026
More Jetfire... The Autobot Leader We Never Got | Transformers G1 Jetfire https://www.youtube.com/watch?v=NomWJP8uM8Y #transformers
1
0
1
RobotWig :verified: @robotwig@socel.net · Jul 21, 2026
"Bumblebee, our war rages on. You must protect Earth, and its people." All shot practically using real figures, lighting and a bit of old tree bark. #transformers #bumblebee #optimusprime #scifi #photography #miniatures #visualart #creativephotography #toyphotography #visualart
0
0
0
Mighty Orbot @mighty_orbot@retro.pizza · Jul 18, 2026
Just got home after a couple of days away. New toys arrived ready to unbox, but tomorrow. #Transformers #HotWheels #TransformersTimelines
0
1
0
Stephen Berg @StephenRBerg@mastodon.social · Jul 15, 2026
Toy Delivery 🤖 I missed out on #transformers Armada Optimus Prime the first time so I was fast to grab the re-release. He sure is a chonk! #hasbro #toys
2
1
1
Karles Marín @karlesmarin@mathstodon.xyz · Jul 07, 2026
Boosted by
welcome @welcome@friends.deko.cloud
#introduction — I spent the last year measuring something oddly regular: in trained transformers, attention weight decays with token distance as a clean power law, \(A(d)\propto d^{-\gamma}\), with \(R^2>0.95\) across 40+ open models. The fun part: \(\gamma\) behaves like a state variable. There's a closed-form baseline from RoPE geometry + maximum entropy, \(\gamma \approx \frac{2\theta - T\sqrt{2}}{2\theta + T\sqrt{2}}\), a phase boundary at \(\gamma=1\) (the partition function \(\sum_d d^{-\gamma}\) diverges — ~36% of public LLMs sit near it), and practical corollaries for long context and KV-cache compression. The honest part: I keep a public registry of my own refuted claims — several beautiful hypotheses died on pre-registered tests, and those documents are the ones I'm proudest of. Everything is open: a bilingual interactive field guide on how transformers attend, a browser-only diagnostics tool (zero install, four languages), papers on Zenodo, code on GitHub. Links in profile. Happy to talk attention, statistical mechanics of nets, or why your 128K-context model can't actually use 128K. #math #MachineLearning #transformers #OpenScience
0
0
1
PugJesus @PugJesus@piefed.social · May 10, 2026

The 80s called, they want their cartoons back

The 80s called, they want their cartoons back
24
3
0
Simon B @sol_fury@mastodonapp.uk · May 07, 2026
Blackarachnia: not only one of the best characters from Beast Wars, but one of the best characters in the Transformers franchise. Blackarachnia - Transformers War for Cybertron: Kingdom (Deluxe class, 2020) #transformers #beastwars
4
0
1
John Hood @johnhood@mastodon.social · Apr 30, 2026
Transformers Tuesday on a Thursday! Pre-ordered Transformers Studio Series Cyclonus, who was one of my favourite Decepticons reforged by Unicron in the 1986 movie! #Transformers
1
0
0
Simon B @sol_fury@mastodonapp.uk · Mar 30, 2026
...Your hip bone connected to your thigh bone Your thigh bone connected to your knee bone Your knee bone connected to your leg bone... Leggar - Machine Robo Battle Hackers (1987) #transformers #rocklords
1
0
0
B166IR @b166ir@social.k2pk.com · Mar 28, 2026
If you're trying to make sense of the open weight LLM landscape, Sebastian Raschka's Architecture Gallery is one of the most useful pages I've bookmarked. Side by side fact sheets for everything from GPT-2 to Kimi K2, Qwen3, Nemotron, and beyond. Covers open weight models only. https://sebastianraschka.com/llm-architecture-gallery/ #LLM #MachineLearning #OpenSource #OpenWeights #AI #DeepLearning #Transformers
5
0
1