Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Abstract
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation. Code and additional visualizations are available on our project page at https://jackswl.github.io/honeycomb/.
Community
We introduce Honeycomb, a video world model built on our proposed HexMemory. HexMemory represents scene features using a low-rank factorization into three spatial and three spatiotemporal planes, whose dimensions remain fixed throughout generation. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency compared to prior spatial-memory methods!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LOCI: Spatial Linear Memory for Streaming World Models (2026)
- WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory (2026)
- Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation (2026)
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation (2026)
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency (2026)
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation (2026)
- GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.37690 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper