My system for the local model is a Strix Halo (Framework Desktop) with 128gb vram. This new Qwen model does not needed to be fully loaded onto VRAM. It can outsource it's ngram embeddings onto normal RAM. But you still need like 90gb of VRAM to run it with good context length I think.