Dude didn't have any quants and I appreciate his approach. Letting KLD go makes sense because we're literally trying to diverge. That's... the point.

So I whipped up a couple. MTP seems to still be working, which is nice. Doesn't do as much on AMD as I wish it did, but it still flies.

BTW, importance Quants are not great for MTP if those tensors get touched. You're better off just leaving them or squanching them down to Q8_0 at most. The importance matrices don't exercise enough of the speculator, it's just gonna gimp it.

Downloads last month
668
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sandmanbuzz/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MTP-gguf

Base model

Qwen/Qwen3.8-27B
Quantized
(29)
this model