Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 12 days ago • 65
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching Paper • 2609.35673 • Published 13 days ago • 34
SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis Paper • 2609.34479 • Published 13 days ago • 30
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published Sep 7 • 124
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published Sep 7 • 124
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published Sep 7 • 124
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens Paper • 2509.03025 • Published Sep 3, 2025
Augmentation-Driven Metric for Balancing Preservation and Modification in Text-Guided Image Editing Paper • 2410.11374 • Published Oct 15, 2024
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias Paper • 2404.00384 • Published Mar 30, 2024