I am using the PrismaAQUA
standard 9.7 t/s
+ Dflash2 30 t/s
+ torch-compile 37 t/s
c8 = 177 t/s
Also, 4-bit has measurable intelligence loss. Sometimes worth it, but, at this size models are barely smart enough at 8 or 6.
The output quality is higher. It's held at full precision (not quantized).
I am using the PrismaAQUA
standard 9.7 t/s
+ Dflash2 30 t/s
+ torch-compile 37 t/s
c8 = 177 t/s