Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published 17 days ago • 69
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 19 days ago • 24
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization Paper • 2608.12314 • Published Aug 12 • 28
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published Jul 14 • 30
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 75
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Paper • 2605.03042 • Published May 4 • 149
GenLIP Collection Model weights of paper "Let ViT Speak: Generative Language-Image Pre-training" • 6 items • Updated May 5 • 8
GenLIP Collection Model weights of paper "Let ViT Speak: Generative Language-Image Pre-training" • 6 items • Updated May 5 • 8
GenLIP Collection Model weights of paper "Let ViT Speak: Generative Language-Image Pre-training" • 6 items • Updated May 5 • 8