arXiv:2609.07498nvidia/nemotron-3.5-lightning-30b-a3bSeptember 7, 2026

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial MovementsExplained for Beginners

Hongxiang Zhao, Mutian Xu, Zeyu Jin +3 more

RoboticsComputer Vision and Pattern Recognition

Abstract

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.

Here is the structured explanation of the CosmoH2G paper.

1. The Problem

Imagine trying to teach a robot to perform a complex task, like flipping a pancake or unscrewing a jar lid, by showing it a video of a human doing it. This is the allure of "Hand-to-Gripper Transfer"—using cheap human demonstrations to train robots. It is cost-effective because we don't need to program every single move; we just need the data.

However, there is a catch. Current methods in this field are like people who have only ever learned to move their hands in flat, two-dimensional gestures. They are great at simple picking-and-placing, but they stumble the moment a task requires a "spatial" movement—a rotation of the wrist, a flip, or a complex trajectory in 3D space. For a robot to be truly useful in a kitchen or a factory, it needs to handle these intricate motions. The paper identifies that existing datasets and methods simply don't have the "complexity budget" to learn these movements reliably.

2. How It Works (The Technical Mechanics)

The authors tackle this by building a massive dataset and a clever two-stage AI architecture.

The Dataset: A Library of "Complex" Hands The first hurdle is data. The team developed a scalable pipeline to capture hand-gripper pairs. They used a "handheld gripper" (a device worn on the hand) to record human motions. Crucially, they filtered for complexity, resulting in 6,189 episodes involving 1,254 unique objects. This is a significant jump in scale and difficulty over previous benchmarks, which often feature simple, repetitive motions.

The Model: The "Keyframe" Strategy The core challenge is that trying to teach a robot to move its gripper from start to finish in one go is unstable. Small errors compound, much like a typo in the middle of a sentence ruins the whole meaning. To fix this, the authors propose a two-stage framework:

  • Stage I (The Director): The AI first predicts only two "keyframes": where the gripper starts and where it ends. Think of this as telling a animator, "The character starts here and finishes there." This simplifies the problem because the AI only needs to figure out the big picture, not every micro-movement in between.
  • Stage II (The Animator): Once the keyframes are set, a second AI generates the smooth, continuous motion in between. Because the start and end points are already "correct," this stage is much easier and more stable.

The "Orientation vs. Translation" Trick The authors noticed that while the gripper needs to rotate (orientation) to match the hand, its position (translation) is more rigid. To prevent the robot from drifting off-course over time, they cleverly decouple the learning:

  • The neural network learns the orientation (which way the fingers point).
  • The translation (exact x, y, z position) is not learned by the network; instead, it is "post-optimized" using simple rules (heuristics) and physics constraints. This acts like a safety net, ensuring the robot doesn't accidentally knock the object over while trying to rotate it.

3. Key Results & Benchmarks

The results show that this two-stage approach is significantly more robust than simply trying to predict the whole path at once.

  • The Metric: In simulations and real robot tests, the framework "significantly outperforms traditional baselines."
  • The Impact: While the paper reports standard success rates, the true takeaway is in the stability. By separating the prediction of keyframes from the generation of the path, the robot makes far fewer "correction" moves. In plain language: the robot is able to complete complex flipping or rotating tasks without dropping the object or jerking wildly, which is a common failure mode in this field.

4. Why It Matters

  • Broader Accessibility: This moves hand-to-gripper transfer from a niche academic curiosity to a viable approach for real-world robotics. If robots can learn complex motions from simple demos, we get cheaper, faster robot deployment.
  • The "Sim-to-Real" Bridge: The rigorous data collection protocol (using the handheld gripper) means the gap between human motion and robot motion is smaller, making it easier to transfer skills from simulation or video to physical robots.
  • What to Watch For: The main limitation is the "keyframe" assumption. If a task requires a movement where the start and end points are similar but the path in between is uniquely critical (like a tight weaving motion), this two-stage approach might need adjustment. Additionally, the reliance on "post-optimizing" translation means the system is highly dependent on the accuracy of the grasping heuristics—if the object is oddly shaped, the "safety net" might not catch everything.

Summary: CosmoH2G is essentially building a massive, complex "library" of human hand movements and teaching robots to sketch the big picture (keyframes) first, then fill in the details, all while double-checking the robot's position to ensure it doesn't drop the object. This makes robots much more capable of handling the twists and turns of everyday manipulation.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →