Elektrine lite

← Feed

@Hexarei@beehaw.org

Post #2479431

2026-03-24 05:25 UTC

> run a local LLM like Claude! > Look inside > "Run ollama" Ollama will almost always be slower than running vllm or llama.cpp, nobody should be suggesting it for anything agentic. On most consumer hardware, the availability of llama.cpp's --cpu-moe flag alone is absurdly good and worth the effort to familiarize yourself with llamacpp instead of ollama.

Replies (1)

  • @Quibblekrust@thelemmy.club 2026-03-25 19:20

    > --cpu-moe ![](https://thelemmy.club/pictrs/image/79e2467d-1bfe-47ec-b918-0a1f4723cf44.jpeg) ::: spoiler AI Acknowledgement The joke is worth the slop, imo. "Cpu Moe". 😂 Find me an anime drawing of a CPU (especially an iconic one) and I'll use that instead. :::

    Open ##2479441