Damus
Laan Tungir · 5d
Imagine what it's going to be like when we can run really capable models locally at hundreds of tokens a second! Reminds me of the early days always imagining having more bandwidth / ram / storage ...
Rod · 5d
I have been thinking of this. There was a startup that made an LLM ASIC for one of the earlier OSS models, I don't remember which, but it was back when those models sucked, so it just sucked faster. Making one for this level will be amazing. Inevitable, IMO