Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction Paper • 2605.26230 • Published May 25 • 38
World Observer: Joint Actor-Observer Generation for Persistent World Modeling Paper • 2610.02162 • Published 5 days ago • 82
EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos Paper • 2609.39378 • Published 6 days ago • 61
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 6 days ago • 116
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering Paper • 2609.38177 • Published 7 days ago • 70
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control Paper • 2608.24115 • Published Aug 25 • 14
Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization Paper • 2606.11180 • Published Jun 9 • 36
WorldKV: Efficient World Memory with World Retrieval and Compression Paper • 2605.22718 • Published May 21 • 41
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Paper • 2506.13654 • Published Jun 16, 2025 • 44
HippoCamp: Benchmarking Contextual Agents on Personal Computers Paper • 2604.01221 • Published Apr 1 • 32
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition Paper • 2602.08439 • Published Feb 9 • 28
Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression Paper • 2512.05081 • Published Dec 4, 2025 • 33
Visual Representation Alignment for Multimodal Large Language Models Paper • 2509.07979 • Published Sep 9, 2025 • 84