OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 11 days ago • 64
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 11 days ago • 64
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Paper • 2609.04298 • Published Sep 9
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design Paper • 2605.02937 • Published Aug 10
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus Paper • 2606.15345 • Published Jun 13 • 17
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos Paper • 2607.00491 • Published Jul 1