its nice to see performance optimizations in the ai space. feels like we will be able to run opus quality models on current home pc(s).
i have seen and tried very cool projects able to fit and run GLM5.2 on my laptop without worrying about memory space with decent speed (decent when you think about what you normally need to run the model)
deepseek also doing a lot advancements in the space. when you combine all in same way, in 5-10 years maybe?