Post #4375253
2026-08-04 08:05 UTC
@Xavier@infosec.exchange I know this. And I disagree about your TCO.
I'm helping a friend with a $2500 computer, that uses approx $90 of electricity a month to run models which are inferior to those that I can rent off Mistral for about $40/month. And those are inferior - though not so much in practice - to frontier models like Opus which are impossible to host yourself, both technically and legally.
When I runf Devstral2 on rented machines (hourly billed GPU VPSes) I can get comparable window sizes to that $40/month hosted version for about $800/month, provided I carefully engineer provisioning that bring the VPS down when unused.
And I haven't then counted the hours tuning, deploying, monitoring etc that. Which could easily be another $1000/month.
I then calculated the $/token or $/month for hosted models that in features/capabilities compare to the one in the $2500 computer at $90/month and it would've cost us under $5/month for the same usage.
In other words: running models that are practically comparable to what's offered at $/token or $/month online currently costs way more than that.
This obv. ignores all the other benefits of running yourself! I only look at the economics here, which, I know and admit, is too narrow.
Replies (1)
-
@Xavier@infosec.exchange 2026-08-04 12:36
@berkes@mastodon.nl I'm sorry, what's your point? I'm building great software and my family has a private AI assistant. And my investment was only $1200 for the 3090 and about $30/day in electricity, which only went up about $2-5 when I upgraded my video card. Home LLMs aren't free, but its not the doomsday scenario you painted.