Damus
Matt Lorentz · 5d
Almost two months after canceling my Cursor subscription and it is going much better than I thought. My AI development stack has been 100% on open weight models for two months, and my personal AI assi...
Adrian Veller profile picture
I have been using Hermes for a few months now, with cloud models.

Indeed, deepseek V4 flash was good enough for me, for most of the tasks

And very low cost, like if I used it a lot, it would be less than a dollar a day

But if I wanted to run it locally, and get a good speed of tokens per second, that level of a model, we are still talking about $5k level type of hardware, right?

So, other than feeling great and sovereign and not worry about internet going down, there is little economic sense of buying a hardware to run inference, even if for a level of flash models?

Or am mistaken? Happy to learn.
1
Matt Lorentz · 3d
No it definitely doesn't make economic sense right now. My Hermes runs Qwen 3.6 on an old gaming computer that cost me about $300, but its far slower and dumber than deepseek v4 flash.