Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I am getting 14 t/s on my 16 GB card at full context with the UD-Q3_K_XL quant. Model link: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.


For some reason the unsloth models leave hardly any room for context. I've switched to the regular (non-unsloth) and get about 25 t/s and get about 80,000 more context tokens for the same quant.


Wow, interesting. What KV cache quantization do you use?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: