Post #1737817
2026-04-12 00:07 UTC
@colincornaby Great points, but consider an alternative: Huge models can be used to distill smaller, focused models. That’s the nut they’re trying to crack. That’s why Intel and AMD and Apple are rushing to add neural cores.
It’s completely untenable economically if the only way they can run useful models is in the datacenter.
Replies (1)
-
@colincornaby@mastodon.social 2026-04-12 00:11
@j_s_j Maybe, not sure what Qwen being a few steps behind Opus means. I think we'll constantly get better at wringing a little bit more out of LLMs. But distillation always comes with some capability loss. I don't know if you distill Mythos if you just end up with Opus again. I think Anthropic said training costs for Mythos were 10 billion, even with distillation base models would be expensive. I tend to agree with the thinking that LLMs are mostly stuck, it's the harnesses that are advanced.