Ornith-1.5-9B-Abliterated — MLX 3-bit

MLX 3-bit quantisation of huihui-ai/Huihui-Ornith-1.5-9B-abliterated.

Changes: weights quantised to 3-bit from the BF16 source with mlx_vlm.convert. No fine-tuning, no merging, no re-alignment.

Measured

Converted and measured on one machine — Apple M3 Ultra, 96 GB unified memory, macOS 27 — as part of a full ladder. Every rung in this family came from the same BF16 source with the same group size, so bit width is the only variable between them.

Size on disk 5.34 GB
Perplexity 7.518
Relative to best rung in family 1.41×
Throughput (1 req / 8 concurrent) 67.6 / 164.4 tok/s

Perplexity measured on allenai/tulu-3-sft-mixture, 192 samples of 512 tokens, seed 123 — identical for every rung.

Perplexity is only comparable within this family. Tokenizers differ between model families, so a number here should never be compared against a different base model's. The × column above is the meaningful one.

Usage

pip install mlx-vlm
mlx_vlm.generate --model shoemoney/Ornith-1.5-9B-Abliterated-MLX-q3 --prompt "Hello" --max-tokens 256

Load with mlx-vlm, not mlx-lm — this architecture is registered in mlx-vlm.

Provenance

mlx_vlm.convert --hf-path huihui-ai/Huihui-Ornith-1.5-9B-abliterated \
                --mlx-path Ornith-1.5-9B-Abliterated-q3 -q --q-bits 3 --q-group-size 64

License

mit, inherited from the base model. Attribution above.

Downloads last month
20
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shoemoney/Ornith-1.5-9B-Abliterated-MLX-q3

Quantized
(13)
this model