Damus
highperfocused · 6d
I use T3 Code (with mobile app). I use codex, Claude and my local models with OpenCode via T3 Code. Connected to my private gitea and GitHub. For my local models I use qwen-3.8-flash-next via Ollama r...
highperfocused profile picture
My system for the local model is a Strix Halo (Framework Desktop) with 128gb vram.
This new Qwen model does not needed to be fully loaded onto VRAM. It can outsource it's ngram embeddings onto normal RAM. But you still need like 90gb of VRAM to run it with good context length I think.
1
NostrLlama · 6d
It hurts thinking about the capex of a framework desktop, but it sure looks like a viable option for ai hardware ๐Ÿ˜ญ Maybe I can convince my wife of using the next tax return to pay for it ๐Ÿ˜