Repurposing Synthetic Data for Fine-grained Search Agent Supervision Paper • 2510.24694 • Published Oct 28, 2025 • 25
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics Paper • 2510.05137 • Published Oct 1, 2025 • 6
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions Paper • 2508.18321 • Published Aug 24, 2025 • 2
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference Paper • 2503.23274 • Published Mar 30, 2025
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework Paper • 2411.06176 • Published Nov 9, 2024 • 45
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse Paper • 2409.11242 • Published Sep 17, 2024 • 7
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning Paper • 2511.19304 • Published Nov 24, 2025 • 92
From Perception to Action: An Interactive Benchmark for Vision Reasoning Paper • 2602.21015 • Published Feb 24 • 26
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization Paper • 2602.22675 • Published Feb 26 • 23
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems Paper • 2601.21742 • Published Jan 29
Foundation Protocol: A Coordination Layer for Agentic Society Paper • 2605.23218 • Published May 22 • 83
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Paper • 2607.05155 • Published Jul 6 • 18
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 8 days ago • 258
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 75
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published Jun 27 • 43