@DaveMWilburn@infosec.exchange
Post #4410468
2026-08-06 04:11 UTC
@Viss@mastodon.social
LLMs are already basically compression.
Discussion in videos from 3blue1brown:
https://youtu.be/l6DKRf-fAAM
https://youtu.be/GlYgs6v2YfU
As far as shrinking the models themselves, Google is investing a lot in this space. They really want to sell you Android phones that can do at least some of the GenAI processing on the device itself, rather than shoveling everything up to a bunch of expensive datacenters. Gemma 4 has a bunch of tweaks that are beyond my understanding to reduce the number of parameters that are loaded in RAM during operation.
Of course, even this approach doesn't work well when there's a global RAM shortage impacting both datacenters and consumer devices. Acquiring even modest amounts of RAM is like buying unobtanium.
Replies (1)
-
@nyanbinary@infosec.exchange 2026-08-06 07:57
@DaveMWilburn@infosec.exchange @Viss@mastodon.social oh neat, we found a way to make the battery life of my phone even shorter :3.