Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 24 days ago • 163
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 23 days ago • 219
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 157
KVpop -- Key-Value Cache Compression with Predictive Online Pruning Paper • 2607.05061 • Published Jul 6 • 22
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting Paper • 2606.27821 • Published Jun 26 • 5
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 52
Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning Paper • 2602.21103 • Published Jun 2 • 9
Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents Paper • 2605.25971 • Published May 25 • 16
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 86
Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring Paper • 2605.16386 • Published May 11 • 2
Leveraging Verifier-Based Reinforcement Learning in Image Editing Paper • 2604.27505 • Published Apr 30 • 28
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 130
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 110
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning Paper • 2604.02721 • Published Apr 3 • 84
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines Paper • 2603.06679 • Published Mar 30 • 5