Elektrine lite

← Feed

@DaveMWilburn@infosec.exchange

Post #4410468

2026-08-06 04:11 UTC

@Viss@mastodon.social LLMs are already basically compression. Discussion in videos from 3blue1brown: https://youtu.be/l6DKRf-fAAM https://youtu.be/GlYgs6v2YfU As far as shrinking the models themselves, Google is investing a lot in this space. They really want to sell you Android phones that can do at least some of the GenAI processing on the device itself, rather than shoveling everything up to a bunch of expensive datacenters. Gemma 4 has a bunch of tweaks that are beyond my understanding to reduce the number of parameters that are loaded in RAM during operation. Of course, even this approach doesn't work well when there's a global RAM shortage impacting both datacenters and consumer devices. Acquiring even modest amounts of RAM is like buying unobtanium.

Replies (1)