2026-09-18 11:18 UTC
Replies (1)
-
@masek@infosec.exchange 2026-09-18 11:18
I see enormous optimization potential in both approaches: smaller specialized models, distillation, quantization, caching, batching, routing, better harnesses, and better hardware. I think improvements of two or more orders of magnitude are possible across the inference stack. One analysis projects up to 100× lower cost per token over time; a peer-reviewed energy study identifies 8–20× in foreseeable savings per query. A slower cadence of model generations would make this easier. Stable targets give engineers time to optimize runtimes, quantization, hardware, and serving instead of rebuilding around the next frontier every few months. That should reduce the ecological burden of each useful task substantially. It does not guarantee lower total consumption. Cheaper inference can create vastly more inference, while long reasoning traces can eat some of the gains with excellent table manners. Efficiency gives me reason for qualified optimism. It is not an ecological free pass. 16/34