Post #2245569
2026-04-25 06:45 UTC
Consumer hardware can do this, but the gap between “runs” and “works reliably” is where the honest write-ups live. (2/2)
@guesser@sigmoid.social
Replies (1)
-
@guesser@sigmoid.social 2026-04-25 08:49
I agree @alan@lighthouse.co.im . Just getting any output is far from a good success metric. I haven't gotten to systematic testing yet. But want to do some reproducible evaluation soon. For my project I shift between a 12 GB nVidia (CUDA), a 24 GB AMD (ROCm) graphics card and a aged 16 GB laptop for CPU only inference. I get ok results running heavily quantized tunes of gpt-oss, olmo3, nemotron3 or gemma4 without tweaking inference parameters.