Damus
NostrLlama · 1w
So much going on lately in the agentic coding bubble. More powerful and smaller local models, private inference via #routstr and decision models like Jev working in parallel... My own system, that I ...
highperfocused profile picture
I use T3 Code (with mobile app). I use codex, Claude and my local models with OpenCode via T3 Code. Connected to my private gitea and GitHub.
For my local models I use qwen-3.8-flash-next via Ollama routed through my LiteLLM Proxy (for statistics and key management).
I got a spare MacBook that runs t3code server and I connect to it via NetBird/Tailscale :)

Works beautiful!
3
highperfocused · 1w
This is how I do most of my coding stuff right now
highperfocused · 1w
My system for the local model is a Strix Halo (Framework Desktop) with 128gb vram. This new Qwen model does not needed to be fully loaded onto VRAM. It can outsource it's ngram embeddings onto normal RAM. But you still need like 90gb of VRAM to run it with good context length I think.
NostrLlama · 1w
Sounds like something I need to explore further. Never heard of T3 Code before