Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 11 days ago • 31
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 11 days ago • 149
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated 9 days ago • 401k • • 856
On-Policy Self-Distillation in Diffusion Models Paper • 2608.24646 • Published 17 days ago • 67
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 18 days ago • 64
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 21 days ago • 63
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published 22 days ago • 111
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published 17 days ago • 70
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 18 days ago • 206