HPC-Quantize / hexstate_requantize.py

Commit History

optimizations for normal llama
ef5ecbc
verified

CompressedGemma commited on

Requires custom llama to use with this quant method
05e9fe6
verified

CompressedGemma commited on

Prototype: Produces NAN values, will not work in llama-server, do not download
145817b
verified

CompressedGemma commited on

Fix OOM on large tensors
92ec61a
verified

CompressedGemma commited on

DC cancellation buff
b57acee
verified

CompressedGemma commited on

I prefer Q2 honestly.. but leaving IQ1 staging in for posterity
3ba66cb
verified

CompressedGemma commited on

Super small quants
4856584
verified

CompressedGemma commited on

Lower RMSE, improve error shaping
bc08931
verified

CompressedGemma commited on

--analog-imatrix: Reduce error in output by 15%
c3e59d2
verified

CompressedGemma commited on

Replace Shor's with accelerated sieve
ae547e6
verified

CompressedGemma commited on

Experimental support for other LLMs
ae8c38d
verified

CompressedGemma commited on

It's only calibrated for Gemma, atm.
07b428c
verified

CompressedGemma commited on