338K clips · 500 household objects · 194 tasks
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
Abstract
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet embodiment-aware teleoperation data remains costly to collect. Egocentric human videos offer a scalable alternative, but human and robotic hands differ profoundly in appearance, articulation, kinematics, and camera-relative geometry. HandEdit formulates this gap as an embodiment-aware image-editing problem. The benchmark asks an editor to replace the visible human hand or hand-arm region with a requested dexterous robot embodiment, while preserving object state, task semantics, contact relationships, viewpoint, and surrounding scene structure.
Embodiment Transfer
Apple cutting
HandEdit Dataset
Articulated objects · tool use · grasping
75 objects · 150 tasks
Room-scale diversity · 54 tasks
Pick-and-place · handover · tool use
Robot Embodiments
Benchmark Results
| Model | Access | LPIPS ROI ↓ | FID ROI ↓ | Removal ↑ | Struct ↑ | ID ↑ | Interaction ↑ | VLM ↑ |
|---|
Conclusion
GPT-Image-2 provides the strongest overall baseline.
VLM-based judgment is useful but not sufficient.
Perceptual quality alone does not guarantee editing task success.
The main challenge lies in embodiment-aware editing.
Qualitative Results
BibTeX
@article{yang2026handedit,
title={{HandEdit}: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing},
author={Zhenjie Yang and Xingyu Jiao and Guopeng Zhong and Shuzhe Yang and Shi Che and Chao Wu and Chenyu Jiang and Dongjie Zhang and Yideng Zhang and Zheng Zhang and Muyun Jiang and Haisheng Su and Shuang Jin and Donghang Zhang and Chao Yang and Li Chen and Hongyang Li and Zuxuan Wu and Yu-Gang Jiang and Xiaosong Jia and Junchi Yan},
year={2026},
eprint={2608.12122},
archivePrefix={arXiv},
primaryClass={cs.RO}
}