Post #1937694
2026-03-23 22:09 UTC
I'm so sick of "local LLM" demos showing some "XXX tokens/second" of output speed but with like zero input context. What's needed for usefulness is like 10k-100k input tokens minimum. All tokens/sec benchmarks should include input context size.
Replies (0)
No replies.