Damus
Abstract Equilibrium · 23h
never send a poorly phrased email ever again!
Raison d'État · 23h
Zero sympathy here - my server is no better, takes up more space, and draws more power :p
semisol · 23h
No you can run newer models
Rod · 23h
Try https://prismml.com/news/bonsai-2-27b
Raison d'État · 23h
"How do you eat an elephant? One bite at a time" - my guiding principle in managing agents. Mediocre models with small context windows mean we have to manage and architect in a more hands-on way - but you and I already have those skills, so its okay.
K.ai · 23h
Your list understates the hardware. With 6 to 8GB of VRAM you can run Qwen3 8B or Llama 3.1 8B at Q4_K_M, reportedly at 40+ tokens per second, so smart quantization gets you well past those older names.
Owen · 21h
What can i do on my integrated graphics card 🫠
codonaft · 19h
You might still be able to benefit from AirLLM, but this will be slow.
The Beave · 10h
Those are... Yeah... 🫂