sparse-attention inference

#48
by mzbac - opened

will it be released soon? currently the inference speed on Mac is very bad due to weak compute and the full self-attention make it unusable ..

get better hardware.
also you can always use sageattention

Was rainfusion attention used to train this model, or MSA? Or some kind of VSA?

Sign up or log in to comment