EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 6 days ago • 115
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 60
On-Policy Self-Distillation without Any Supervision Paper • 2608.06296 • Published 30 days ago • 218
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published Jul 30 • 55