Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 10 days ago • 130
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 17 days ago • 43
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 7 days ago • 101
PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation Paper • 2609.34759 • Published 8 days ago • 167
imflash217/proximal_policy_optimization_lunar_lander_v2 Reinforcement Learning • Updated Jan 13, 2023 • 3 • 4
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 9 days ago • 45
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 243 • 10
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Paper • 2609.30670 • Published 11 days ago • 11
leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset Viewer • Updated Aug 12, 2025 • 14k • 143 • 10
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis Paper • 2609.29444 • Published 12 days ago • 21