Post #2589529
2026-05-08 01:56 UTC
You can use something like KoboldCPP on Linux, which allows both RAM and VRAM combined to run a model. O’course, not as fast when compared to pure VRAM or the Mac approach, but it is an option. I use my 128gb RAM with some GPUs for running models.
Replies (1)
-
@boonhet@sopuli.xyz 2026-05-08 09:46
Ollama and llama.cpp allow it too but it’s super slow in my experience.