Hitch · 4w Im getting 50tk/s with qwen3.6 35B FP8 on vllm with speculative decoding. 250k tk context window. Have to do a lot of configs to get things working right ABH3PO @ABH3PO 1781926609 Thats considerably lower than what I would have thought especially for a 35b model, do >70b models run at similar speeds? 1🧡1