Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Abstract
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Community
Why do some rewards get learned much more slowly in multi-reward RL, even with GDPO? We show the root cause is advantage energy: under GDPO, a reward's energy scales with how often it is active in a batch, so sparse rewards are drowned out. DARA fixes this with a density-based weight that, in theory, gives every reward equal energy. It is recomputed from each batch and needs no tuning. It reaches the target behavior in up to 26% fewer steps on tool calling and 65% fewer steps on math, with competitive final performance.
Code: github.com/zhaihaotian/DARA
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (2026)
- ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients (2026)
- SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation (2026)
- CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning (2026)
- SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning (2026)
- Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR (2026)
- Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.00574 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper