338K clips · 500 household objects · 194 tasks
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
Accepted by NeurIPS 2026
Abstract
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet embodiment-aware teleoperation data remains costly to collect. Egocentric human videos offer a scalable alternative, but human and robotic hands differ profoundly in appearance, articulation, kinematics, and camera-relative geometry. HandEdit formulates this gap as an embodiment-aware image-editing problem. The benchmark asks an editor to replace the visible human hand or hand-arm region with a requested dexterous robot embodiment, while preserving object state, task semantics, contact relationships, viewpoint, and surrounding scene structure.
Embodiment Transfer
Apple cutting
HandEdit Dataset
Articulated objects · tool use · grasping
75 objects · 150 tasks
Room-scale diversity · 54 tasks
Pick-and-place · handover · tool use
Robot Embodiments
Benchmark Results
| Model | Access | LPIPS ROI ↓ | FID ROI ↓ | Removal ↑ | Struct ↑ | ID ↑ | Interaction ↑ | VLM ↑ |
|---|
Conclusion
GPT-Image-2 provides the strongest overall baseline.
VLM-based judgment is useful but not sufficient.
Perceptual quality alone does not guarantee editing task success.
The main challenge lies in embodiment-aware editing.
Qualitative Results
BibTeX
@article{yang2026handedit,
title={HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing},
author={Yang, Zhenjie and Jiao, Xingyu and Zhong, Guopeng and Yang, Shuzhe and Che, Shi and Wu, Chao and Jiang, Chenyu and Zhang, Dongjie and Zhang, Yideng and Zhang, Zheng and others},
journal={arXiv preprint arXiv:2608.12122},
year={2026}
}