Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
Abstract
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
Community
On Blackwell B200, tensor cores outrun the exponential unit by ~500× (8192 vs 16 ops/clk/SM), so exp becomes exposed in fused attention. We ask what a frozen pretrained LLM actually needs from softmax — and use the answer to replace exp2 inside FlashAttention-4.
What frozen models need (10 models, 0.5B–72B, no retraining):
- Attention support and within-row resolution can be cut substantially — but uniform weighting on the same support hurts at every tested layer.
- Where resolution goes matters as much as how much: finer intervals near the row maximum lower NLL in all 10 models, even when overall approximation error goes up.
- A scalar distortion budget is not enough: flattening vs. sharpening at equal attention JSD gives opposite-sign losses depending on the model.
Rowmax-H15 in FA4: weights snap to {1, 1.5}×2^k anchored at the kernel's running row max — no calibration, one FP32 add + one bit shift per element.
- FP8 attention forward on B200: +12.4% (causal 8K), +25.8% (non-causal 8K); 9.3% faster than the best stock FA4 emulation setting we tested
- −8.4% board energy per forward (causal 16K)
- Perplexity +0.09–0.49% across 5 models from 3 families (BF16 kernel, 2K)
Scope: attention forward on B200; Approximate Softmax Characterizing.
Companion training-side paper (pretraining from scratch with quantized softmax — why the backward rule and calibration gradients matter): https://arxiv.org/abs/2609.33591
Kernel patch and code will be released. Questions and feedback welcome!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Softmax Reparameterization for Output-Head Quantization (2026)
- VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention (2026)
- Beyond Sparse Weights: When Is Attention Compressible? (2026)
- Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers (2026)
- VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models (2026)
- Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference (2026)
- SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.33586 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper