Image-to-Video
Safetensors
English
video-generation
text-to-video
memory

MosaiChunk

Learned memory-router checkpoints for MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation.

Project page · Paper · Code · RememBench

Checkpoint Backbone
t2v/model.safetensors RAVEN-adapted MiniMax-H3 (H3-AR)
i2v/model.safetensors LingBot-World-Infinity

Each folder contains router weights and config.json. The weights are exported without changing tensor values; optimizer and training state are excluded.

These checkpoints require their corresponding frozen video backbone. T2V also requires the pretrained RAVEN streaming adapter. Backbone and adapter weights are not bundled here.

Download

from huggingface_hub import snapshot_download

snapshot_download("mosaichunk/MosaiChunk", local_dir="checkpoints/MosaiChunk")

See the code repository for training and inference.

License

Each router checkpoint is subject to the license of its backbone:

The checkpoints do not include backbone or adapter weights and grant no rights to them.

Citation

@misc{zhang2026mosaichunkcompositingspatiotemporalmemory,
      title={MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation},
      author={Yiwen Zhang and Haocheng Xi and Michael Tian-Yue Liu and Alexei A. Efros and Hadar Averbuch-Elor and Qianqian Wang and Haiwen Feng},
      year={2026},
      eprint={2610.02153},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02153},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train mosaichunk/MosaiChunk

Paper for mosaichunk/MosaiChunk