Damus
Hitch · 4w
Im getting 50tk/s with qwen3.6 35B FP8 on vllm with speculative decoding. 250k tk context window. Have to do a lot of configs to get things working right