Post #4505512
2026-07-23 15:49 UTC
@jcoglan@mastodon.social Yes they did - the K/V cache. And prices are cheaper if your subsequent queries with the same prefix are within time that that stays in memory on the server.
Agreed it is extremely inefficient, and mostly LLMs are better if you do controlled prompts with JSON output, rather than long conversation threads. But people are pushed into the long threads by the standard harnesses. That is a waste!
Replies (0)
No replies.