Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents
Abstract
A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.
Community
.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation (2026)
- EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation (2026)
- AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation (2026)
- Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation (2026)
- WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning (2026)
- Recursive Video In-Context Learning for Agentic Robot (2026)
- MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.07785 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
willsuh/attacca-dataset
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper