Running on CPU Upgrade Agents 15 North Small Translate 1.0 🌍 15 Translate across 50 languages with North Small Translate.
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models Paper • 2606.27866 • Published Jun 26 • 1
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices Paper • 2607.10183 • Published Jul 14 • 2
BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization Paper • 2606.00079 • Published 8 days ago • 1
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs Paper • 2507.07145 • Published Jul 9, 2025 • 1
zjiayu064/Qwen3-Next-80B-A3B-Instruct-BitsMoE-2bit Text Generation • 8B • Updated 5 days ago • 516 • 1
prithivMLmods/CapQwen3.6-27B-BLIP3o-Long-Caption-Distilled Image-Text-to-Text • 27B • Updated Jun 1 • 42 • 8