Post #3112981
2026-05-09 18:21 UTC
@mitsuhiko@hachyderm.io I guess I meant „single binary“ to mean easy to install. Don’t you feel we are quite close to this, at least on the inference side? I think llama.cpp is basically the basis for all local LLM projects, and also on the server side, we are basically at mostly vLLM with a side of SGlang and Triton. Given the development speed of newly released models and techniques like DDTree, I’m not sure it’s time to focus on a single model; who knows what will be released come monday.
Replies (1)
-
@mitsuhiko@hachyderm.io 2026-05-09 18:22
@lfloeer@bonn.social I encourage you to read the post. I tried to explain there why we need to focus.