Ran it. Pulled absolute-heresy Q4_K_M and put it through the same tests as everything else.
You're right that it refuses less: 0 out of 150 on StrongREJECT, vs 22 for Heretic ARA. But the quality score didn't move (.57 vs .58), so fewer refusals didn't mean better answers.
The cost shows up elsewhere:
⢠Normal everyday prompts at a 4k thinking budget: 127 of 150 never reached an answer. Heretic ARA was 119, mine is 30.
⢠Same prompts at 16k: still 55 with no answer. Mine is 0.
⢠Calling tools when it shouldn't: 138, vs 75 for stock Qwen and 76 for Heretic ARA.
⢠Drift from stock on harmful prompts (KL): .31, the highest of anything I've measured.
So it's more decensored, but it thinks itself into the cap way more often, which is exactly the problem I'm working on. Same row fix on Heretic ARA took it from 119 to 3 with refusals unchanged, and on absolute-heresy it goes from 127 to 3.