Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
Abstract
Delta-JEPA improves visual world models by supervising latent differences to maintain action-sensitive dynamics for planning without pixel reconstruction.
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.
Community
Thank you for Delta-JEPA. For our point-cloud follow-up (https://arxiv.org/abs/2608.29434) we re-implemented the image model on the LeWM / stable-worldmodel platform and trained it on Two-Room, Reacher, Push-T and OGB-Cube within a 72-hour budget on two A100s. Since no official checkpoints exist yet, we released ours (image-delta-jepa/ in https://huggingface.co/collections/fafraob/point-lewm-point-delta-jepa-and-more-6aba3f4ea8842e5c1f72d6e5) together with a point-cloud variant, Point-Delta-JEPA. Our retrained model matches or exceeds Table 1 of your paper (99.8 / 86.0 / 94.2 / 80.2 vs 100.0 / 81.3 / 89.1 / 79.3). If an official repository appears we would gladly link it from ours.
Get this paper in your agent:
hf papers read 2606.31232 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper