Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
Casteil
15 days ago
|
parent
|
context
|
favorite
| on:
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.
qwen3.5:122b-a10b is significantly faster at around 60-65.
smcleod
15 days ago
|
next
[–]
No magic, just oMLX with MTP. You can look through the speed the community is getting here:
https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...
syntaxing
15 days ago
|
prev
[–]
With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
Casteil
15 days ago
|
parent
[–]
It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
qwen3.5:122b-a10b is significantly faster at around 60-65.