Elektrine lite

← Feed

@florenciocano@infosec.exchange

Post #4060447

2026-07-24 10:31 UTC

Gemma 4 31B quantized is hands down the best model I have been able to run locally on a RTX 4090. I run it with `llama.cpp/build/bin/llama-server -m ~/models/gemma-4-31B-it-Q4_K_M.gguf -c 65536 -ngl 99 -ctk q4_0 -ctv q4_0 -fa on --host 0.0.0.0 --port 8081` Any suggestion to run it more efficiently? Any other model I should try?

Replies (0)

No replies.