Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
Abstract
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
Community
Long-horizon agents need to keep track of the current world state to guide their decisions. Yet maintaining a coherent belief alone does not ensure progress: agents can keep acting without meaningfully advancing their goals, a failure mode this paper calls Belief Trapping. This paper introduces PoS (Progression of States), an inference-time framework that combines explicit belief maintenance with trapping detection, diagnosis, and recovery, without additional training. PoS represents the current world and unresolved information and task requirements, validates belief updates for consistency and evidential support, and monitors whether interactions advance the task. When progress stalls, it diagnoses both the trapping pattern—stagnation, cycles, or drift—and the blocked requirement, then composes targeted recovery constraints to guide subsequent actions. Across four benchmarks covering task execution and evidence-seeking diagnosis, PoS achieves the highest overall performance among the evaluated methods in all 12 benchmark–backbone combinations, with relative gains over the strongest same-backbone baseline reaching 22.68% on ALFWorld and 37.89% on RCA-100 joint accuracy. These results highlight the value of continually maintaining beliefs and restoring meaningful task progress for reliable long-horizon agents.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SKILL.state: Scalable Long-Horizon Agent Skills (2026)
- HarnessWAM: Bridging Prediction and Deliberation in World Action Models (2026)
- Agent-Editing World Model: Rethinking World Modeling for LLM Agents (2026)
- Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments (2026)
- Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents (2026)
- FlowState: Execution State as Memory for Long-Horizon LLM Agents (2026)
- Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper