Memorize When Needed:
Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

1The Hong Kong Polytechnic University
2OPPO Research Institute
*Equal contribution. Corresponding author.
Method overview teaser figure

Abstract

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent content during scene revisits and diminished generative capacity when exploring novel regions, even when trained on extensive annotated data. To address these limitations, we propose a decoupled framework that separates memory conditioning from generation. Our approach significantly reduces training costs while simultaneously enhancing spatial consistency and preserving the generative capacity for novel scene exploration. Specifically, we employ a lightweight, independent memory branch to learn precise spatial consistency from historical observation. We first introduce a hybrid memory representation to capture complementary temporal and spatial cues from generated frames, then leverage a per-frame cross-attention mechanism to ensure each frame is conditioned exclusively on the most spatially relevant historical information, which is injected into the generative model to ensure spatial consistency. When generating new scenes, a camera-aware gating mechanism is proposed to mediate the interaction between memory and generation modules, enabling memory conditioning only when meaningful historical references exist. Compared with the existing method, our method is highly data-efficient, yet the experiments demonstrate that our approach achieves state-of-the-art performance in terms of both visual quality and spatial consistency.

Experimental Results

Loading prompt...
Loading prompt...
Loading prompt...
Loading prompt...
Loading prompt...
Loading prompt...

Results Gallery

Qualitative Comparisons

RealEstate10K Indoor

DFoT

VMem

WorldPlay

Ours

RealEstate10K Outdoor

DFoT

VMem

WorldPlay

Ours

OOD Scene

DFoT

VMem

WorldPlay

Ours

BibTeX

@article{guo2026memorize,
  title={Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation},
  author={Guo, Yanjun and Zhang, Zhengqiang and Wang, Pengfei and Liang, Xinyue and Ma, Zhiyuan and Zhang, Lei},
  journal={arXiv preprint arXiv:2604.18215},
  year={2026}
}

Acknowledgements

If you are interested in our work, please also check out the following related works. We would like to thank the contributors to:
Wan: Open and Advanced Large-Scale Video Generative Models
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
WorldMem: Long-term Consistent World Simulation with Memory
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
SeVA: Stable Virtual Camera: Generative View Synthesis with Diffusion Models
DFoT: History-Guided Video Diffusion
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
VideoX-Fun https://github.com/aigc-apps/VideoX-Fun