OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction Paper • 2610.01762 • Published 5 days ago • 220
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Paper • 2609.36380 • Published 8 days ago • 136
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Paper • 2609.12641 • Published 25 days ago • 71
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published Aug 27 • 84
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning Paper • 2609.03199 • Published Sep 2 • 31
Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training Paper • 2608.24680 • Published Aug 25 • 11
Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models Paper • 2606.16700 • Published Jun 15 • 15
Guava: An Effective and Universal Harness for Embodied Manipulation Paper • 2606.18363 • Published Jun 16 • 28
MolmoAct2: Action Reasoning Models for Real-world Deployment Paper • 2605.02881 • Published May 4 • 358
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents Paper • 2604.18543 • Published Apr 20 • 31
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Paper • 2604.12887 • Published Apr 14 • 5
Emergent Social Intelligence Risks in Generative Multi-Agent Systems Paper • 2603.27771 • Published Mar 29 • 50
Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration Paper • 2603.24800 • Published Mar 25 • 69
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics Paper • 2603.14375 • Published Mar 15 • 19