Elektrine lite

← Feed

@Lydie@tech.lgbt

Post #3955084

2026-07-13 20:04 UTC

@ben@social.benjaminturner.me Nice. Amazing how well these old platforms work for this. Just using them to glue a bunch of GPUs together. What size models does that run and how many tokens per second? I'm getting around 64t/s on my old 2080tis

Replies (2)

  • @Lydie@tech.lgbt I haven't checked the t/s, I'll report back. I've been going back and forth between gemma4:12b with 128k context and qwen3.6-27b. largest I've tried is 35b. Nice part about the 5060 Ti is that it's an 8x card which is what my board supports with two cards installed so it doesn't feel like I'm leaving any performance on the floor (if it even makes a noticeable difference in the first place)

    Open ##3955086

  • @Lydie@tech.lgbt gemma4:12b is around 41t/s, qwen3.6:35b was 97t/s

    Open ##3955087