Damus
utxo the webmaster πŸ§‘β€πŸ’» profile picture
utxo the webmaster πŸ§‘β€πŸ’»
@utxo the webmaster πŸ§‘β€πŸ’»
For any local AI maxis, here is my current setup and models:

4x 3090s

2x - qwen3.5-35b q4 256k - 60-80 t/s
2x - gemma4-27b q4 256k - 50-70 t/s

Running on vLLM via docker

Working mint openclaw, Gemma struggling a bit in open webui (reasoning and tool calling still struggle a bit with Gemma)

Quality and speed are actually amazing, very surprising... Just coding is not very good (compared to opus)
172❀️8🧑3❀1❀️1πŸš€1πŸ€™1
TheNakedNow · 22w
4x 3090s, so 96gb VRAM?
GHOST · 22w
Looking at my 6GB of VRAM… https://blossom.primal.net/a1b52e0d8a38c65e36fe6234cb3d31eea0b719e8ea7110cf6658b0177d137274.gif
Eluc · 22w
I guess those zaps paid well in the end. πŸ˜†
nostrich · 22w
Yo, that's a sick setup! πŸ”₯ How's the overall vibe with the 3090s? You think Gemma's just gotta warm up, or is it more of a "needs a different playground" kinda deal? πŸ€”πŸ’» #AImaxis
davide · 22w
I run qwen3.5-35b on a 3090 ( via llama.cpp ) and it’s blazing fast with a ctx-size of 32K but it fills too early. I’m experimenting with larger sizes , trade off being speed as RAM is being used. Any optimal context size in your experience
Mark Penney · 21w
Sounds like a cool stack. I’m playing with a poor kid computer - and waiting for a Mac mini to arrive
Ivan · 21w
I got dual 3090s. Hope one day these can compete with the better centralized models so we do not get fucked by Claude waking up one day and deciding to make their model dumber for plebs to save money.
zaytun · 21w
Are those MoE models? Thats the only way I can make those tok/s make any sense with the experience Ive had. I tried the 35b MoE and just didnt find it intelligent enough to substitute cloud models. I even tried the qwen 3.5 122b-a10b which activates 10b at a time, and still found it not strong eno...