Elektrine lite

← Feed

@maikel@vmst.io

Post #3106821

2026-06-03 07:47 UTC

Not long ago I got tired of Qwen3.6b-27b at q3 being weird at times and decided I wanted it at larger quant. I found this GPU at a ridiculous price in Bezosland, so went for it. Had to go through the complicated mess of compiling llama.cpp with "-DGGML_BACKEND_DL=ON" and other flags to ensure Rocm and CUDA both were loaded depending on card, so that my Nvidia 5060 with 16GB and that 9060 with 16GB operated as one uniform thing constrained by the speed of my PCIE 3x16 (my mobo is sligthly old). It worked. But then I tried my favourite toy: ComfyUI. And it uses PyTorch underneath it all, guess what? It's.... 1/x #locallm

Replies (0)

No replies.