Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 7 days ago • 6
Softmax Reparameterization for Output-Head Quantization Paper • 2609.31291 • Published 9 days ago • 6
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 7 days ago • 4
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 7 days ago • 4
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 7 days ago • 6
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 7 days ago • 4
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 7 days ago • 6