SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published 8 days ago • 39
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Paper • 2605.15565 • Published May 15 • 17
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Paper • 2604.02029 • Published Apr 2 • 153
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis Paper • 2603.20278 • Published Mar 17 • 102
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65
RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI Paper • 2602.07837 • Published Feb 8 • 57
WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning Paper • 2602.04634 • Published Feb 4 • 100
MARS: Reinforcing Multi-Agent Reasoning of LLMs through Self-Play in Strategic Games Paper • 2510.15414 • Published Oct 17, 2025 • 1