Damus
slembcke · 2w
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpq0q8eanytqf3m9jd0fqvxu6r63lt6jle9qgyxfrd97zv2nwfa8cpq30622w I will say some of the bigger local models like the 31b param Gemma 4 fun as toys for ...
Bartosz Taudul profile picture
@nprofile1q... @nprofile1q... You say you have 80+ GB of RAM available, but run models in the 30B class?

What you want is a ~300B model, like https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF which can be quantized to fit into the 80+ GB you have. This model also has super tiny KV cache size, 1M tokens at q8 is about 3 GB.

Model size >>>> quantization loss.