Elektrine lite

← Feed

@mitsuhiko@hachyderm.io

Post #3112980

2026-05-09 18:06 UTC

@lfloeer@bonn.social I don’t care about single binary :) I need efforts to focus so that one model in one harness in one inference engine works as good as possible. One can go wider from there.

Replies (1)

  • @lfloeer@bonn.social 2026-05-09 18:21

    @mitsuhiko@hachyderm.io I guess I meant „single binary“ to mean easy to install. Don’t you feel we are quite close to this, at least on the inference side? I think llama.cpp is basically the basis for all local LLM projects, and also on the server side, we are basically at mostly vLLM with a side of SGlang and Triton. Given the development speed of newly released models and techniques like DDTree, I’m not sure it’s time to focus on a single model; who knows what will be released come monday.

    Open ##3112981