Request for lightweight Q3 and Q2 GGUF quantizations

#2
by wsq194 - opened

Hi, thank you for releasing this model.

Would you consider adding lightweight GGUF quantizations such as Q3 and Q2 variants? The 27B model is difficult to run on machines with limited VRAM or system RAM, and lower-bit versions would make local testing and inference much more accessible.

If official Q3/Q2 files are not planned, could you recommend a compatible quantization workflow or specific quant types for this model?

Thank you!

Sign up or log in to comment