BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 13 days ago • 30
jaehyeokdoo2/openpi-droid-pnpcarrot-singetask-qflow-offlinerl-criticwarmup2000-alpha100-bs8-test Updated Mar 4 • 3
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts Paper • 2609.24058 • Published 8 days ago • 55
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 10 days ago • 17
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 9 days ago • 50
open-source-metrics/reinforcement-learning-checkpoint-downloads Viewer • Updated Oct 6, 2022 • 367 • 88 • 4
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 11 days ago • 35
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 11 days ago • 136
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 15 days ago • 156
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 12 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 12 days ago • 44
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 12 days ago • 57
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 13 days ago • 37