Elektrine lite

← Feed

@brucethemoose@lemmy.world

Post #3532493

2026-06-25 04:31 UTC

CPU offloading is too slow unless you use a hybrid MoE model, with the --n-cpu-moe parameter, specifically. This only offloads “sparse” parts of the model to the CPU, which take up a lot of RAM but are very compute-lite to run.

Replies (0)

No replies.