Video2Skill: From Streaming Experience to Reusable Embodied Skills
Abstract
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
Community
๐ค Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from?
๐ผ Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention.
๐ง VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library?
๐งต Video2Skill: From Streaming Experience to Reusable Embodied Skills
Why it matters:
๐ It can label much more data for training agents.
๐ It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new?
๐ Paper: https://arxiv.org/abs/2609.36691
๐ Project: https://andyzworks.github.io/video2skill/
๐ค Data: https://huggingface.co/datasets/Sterzhang/video2skill-bench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs (2026)
- RoboChrono: A Real Robot Benchmark for Streaming Task Understanding (2026)
- SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation (2026)
- RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation (2026)
- Uruqi: Learning Spatial Cognition from Visual Experience (2026)
- RoboBridge: A Self-Evolving Embodied Agent Framework for Sim-to-Real Transfer (2026)
- ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36691 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper