Post #2479431
2026-03-24 05:25 UTC
> run a local LLM like Claude!
> Look inside
> "Run ollama"
Ollama will almost always be slower than running vllm or llama.cpp, nobody should be suggesting it for anything agentic. On most consumer hardware, the availability of llama.cpp's --cpu-moe flag alone is absurdly good and worth the effort to familiarize yourself with llamacpp instead of ollama.
Replies (1)
-
@Quibblekrust@thelemmy.club 2026-03-25 19:20
> --cpu-moe  ::: spoiler AI Acknowledgement The joke is worth the slop, imo. "Cpu Moe". 😂 Find me an anime drawing of a CPU (especially an iconic one) and I'll use that instead. :::